Method for Image Processing and Image Processing Apparatus

By processing multiple transnasal endoscopic images captured under different light sources through image registration and segmentation, the method improves segmentation and object identification accuracy in transnasal endoscopy.

JP7711764B2Active Publication Date: 2025-07-23NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023558635
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-24
Publication Date
2025-07-23
Estimated Expiration
2041-03-24

AI Technical Summary

Technical Problem

Existing image processing methods for transnasal endoscopic images captured under different light sources fail to utilize the varying characteristics of lesions or tumors effectively, leading to reduced accuracy in segmentation and object identification due to slight distortions and the lack of integration of information from multiple light source images.

Method used

A method and apparatus that processes a plurality of images captured under different light sources, utilizing image registration and segmentation to synthesize results, improving accuracy by combining segmentation labels and identifying objects based on registered images.

Benefits of technology

Enhances the accuracy of image segmentation and object identification by integrating information from multiple light source images, thereby improving diagnostic reliability in transnasal endoscopy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711764000025
    Figure 0007711764000025
  • Figure 0007711764000026
    Figure 0007711764000026
  • Figure 0007711764000027
    Figure 0007711764000027
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to a method, an apparatus, and a computer-readable medium for image processing. In some embodiments, a method of image processing is provided. The method includes obtaining a plurality of images of an object captured under different light sources, the plurality of images including a target image and at least one associated image. The method further includes generating a segmentation label for the target image based on segmentation results of the plurality of images. In other embodiments, another method, a corresponding apparatus, a computer-readable medium, and a computer program product are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of image processing, and more particularly to methods, apparatuses, and computer-readable media for image processing.

Background Art

[0002] In actual clinical settings, accurate diagnosis requires considering all aspects of the evidence data. For example, in transnasal endoscopic diagnosis, doctors often manually switch different light sources at the same camera position to examine suspicious lesions or tumors. Lesions or tumors under different light sources may exhibit different characteristics, providing a large amount of useful information for doctors to make more accurate and reliable judgments. Therefore, it is reasonable to assume that these characteristics also contribute to improving the performance of image processing tasks (such as semantic segmentation, instance segmentation, and / or object recognition) for transnasal endoscopic images.

Summary of the Invention

Problems to be Solved by the Invention

[0003] Generally, exemplary embodiments of the present disclosure provide a method, an apparatus, and a computer-readable medium for image processing.

Means for Solving the Problems

[0004] In a first aspect, a method for image processing is provided. The method includes obtaining a plurality of images related to an object captured under different light sources, the plurality of images including a target image and at least one related image, and generating a segmentation label for the target image based on the segmentation results of the plurality of images.

[0005] In a second aspect, a method of image processing is provided. The method includes obtaining a plurality of images of an object captured under different light sources, registering the plurality of images, and identifying the object based on the registered plurality of images.

[0006] In a third aspect, an apparatus for image processing is provided. The apparatus includes at least one processor, and the at least one processor is configured to obtain a plurality of images of an object captured under different light sources, the plurality of images including a target image and at least one associated image, and generate a segmentation label for the target image based on the segmentation results of the plurality of images.

[0007] In a fourth aspect, an apparatus for image processing is provided. The apparatus includes at least one processor, and the at least one processor is configured to obtain a plurality of images of an object captured under different light sources, register the plurality of images, and identify the object based on the registered plurality of images.

[0008] In a fifth aspect, a computer-readable storage medium storing instructions is provided. When the instructions are executed on at least one processor, the at least one processor is caused to execute the method according to the first or second aspect of the present disclosure.

[0009] In a sixth aspect, a computer program product including machine-executable instructions is provided. When the machine-executable instructions are executed on at least one processor, the at least one processor is caused to execute the method according to the first or second aspect of the present disclosure.

[0010] It should be understood that the summary of the invention is not intended to identify the important or essential features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present embodiments should be readily understandable from the following description.

Brief Description of the Drawings

[0011] By further describing some embodiments of the present disclosure in the drawings in more detail, the above-mentioned and other objects, features, and advantages of the present disclosure will become more apparent. In the drawings, the same reference numerals generally refer to the same components within the embodiments of the present disclosure.

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0012] Here, the principles of the present disclosure will be described with reference to several exemplary embodiments. It should be understood that these embodiments are described for illustrative purposes only and are intended to assist those skilled in the art in understanding and implementing the present disclosure, without suggesting any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways different from the methods described below.

[0013] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0014] As used herein, the singular forms "a", "an", and "the" include the plural forms as well, unless the context clearly dictates otherwise. The terms "comprising" and its variants should be understood as open-ended terms meaning "including, but not limited to". The term "based on" should be understood as "at least partially based on". The terms "one embodiment" and "embodiment" should be understood as "at least one embodiment". The term "another embodiment" should be understood as "at least one another embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included hereinafter.

[0015] In some instances, values, procedures, or devices are referred to as "best", "lowest", "highest", "minimum", "maximum", etc. Such descriptions are intended to indicate that a selection can be made from among a number of available functional alternatives, and it will be understood that such a selection need not be better, smaller, higher, or otherwise more preferred than other selections.

[0016] As used herein, "neural network" or "network" can process inputs and provide corresponding outputs, and typically includes an input layer, an output layer, and one or more hidden layers between the input layer and the output layer. A neural network usually includes a number of layers connected in sequence, where the output of the previous layer is provided as the input of the next layer, the input layer receives the input of the neural network, and the output of the output layer functions as the final output of the neural network. Each layer of the neural network includes one or more nodes (also referred to as processing nodes or neurons), and each node processes the input from the previous layer. Hereinafter, the terms "neural network", "model", "network" and "neural network model" can be used interchangeably.

[0017] As described above, images of the same object captured under different light sources show different features and can provide a large amount of information useful for improving the performance of image processing. According to conventional image processing solutions, image segmentation (e.g., semantic segmentation, instance segmentation) or object identification can be performed on a single image to obtain a processing result. However, in conventional solutions, these features and useful information cannot be used to improve the performance of image processing. Furthermore, since images of the same object can be captured at different times under different light sources, the objects in the images may have slight distortions. In this case, directly performing image segmentation or object identification on the images may deteriorate the segmentation / identification accuracy.

[0018] Embodiments of the present disclosure provide solutions for image processing to solve the above problems and one or more other potential problems. In some embodiments, a plurality of images regarding an object captured under different light sources can be obtained, where the plurality of images includes a target image and at least one related image. Based on the segmentation results of the plurality of images, a segmentation label for the target image may be generated. In this way, by synthesizing the segmentation results of the images regarding the object captured under different light sources to obtain the final segmentation result for the target image, the segmentation accuracy of the target image can be improved. In some other embodiments, a plurality of images regarding an object captured under different light sources may be obtained and registered. An object may be identified based on the registered plurality of images. In this way, by removing the influence of slight deformations of the object between different images through image registration, the accuracy of object identification can be improved.

[0019] Hereinafter, some exemplary embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be easily understood by those skilled in the art that the detailed description given herein regarding these drawings is provided for illustrative purposes only and does not imply any limitation on the scope of the present disclosure.

[0020] FIG. 1 is a diagram showing an exemplary image processing system 100 capable of implementing an embodiment of the present invention. As shown in FIG. 1, the system 100 may include an image collection device 110 and an image processing device 120. In some embodiments, the devices 110 and 120 may be implemented in different physical devices respectively. Alternatively, the devices 110 and 120 may be implemented in the same physical device. It should be understood that the configuration of the system 100 is shown for illustrative purposes only and does not imply any limitation on the scope of the present disclosure. Embodiments of the present disclosure may also be applied to other systems having different configurations.

[0021] The image acquisition device 110 may collect the image 101 to be processed by the image processing device 120. In some embodiments, the image acquisition device 110 may collect a plurality of images regarding an object captured under different light sources at the same camera position. For example, the different light sources may be associated with different wavelengths or different combinations of wavelengths. In some embodiments, the image acquisition device 110 may be a medical assistance device or an endoscope assistance device. The image 101 may be a medical image such as an endoscope image. For example, as described above, in nasal endoscopy diagnosis, a doctor may manually switch different light sources at the same camera position to examine a suspicious lesion and capture a plurality of images regarding the nasal lesion. The image 101 collected by the image acquisition device 110 may be provided to the image processing device 120. The image processing device 120 may process the image 101 to generate an image processing result 102.

[0022] In some embodiments, the image processing device 120 may perform an image segmentation task. For example, a plurality of images regarding the same object may include a target image and at least one related image. The image processing device 120 may generate a segmentation label for the target image based on the segmentation results of the plurality of images. The image processing result 102 may indicate the segmentation label for the target image. As used herein, the "segmentation result" of an image may indicate the respective probabilities that each pixel in the image belongs to different predetermined categories. For example, the segmentation result may be represented as a heat map in which the brightness of the pixel is used to indicate the probability that the pixel belongs to a certain category. The "segmentation label" of an image may indicate that each pixel in the image belongs to one of the predetermined categories. For example, the segmentation label may be represented as a vector or an array indicating the corresponding category of each pixel, or may be represented as a visual image in which pixels of different categories are identified with different colors.

[0023] In semantic segmentation, the segmentation result of an image may indicate the respective probabilities that each pixel in the image belongs to different predetermined semantic categories. The segmentation label of the image may indicate that each pixel in the image belongs to one of the predetermined semantic categories. Examples of semantic categories may include, but are not limited to, background, person, animal, vehicle, etc. In instance segmentation, the segmentation result of an image may indicate the respective probabilities that each pixel in the image belongs to different predetermined instance categories. The segmentation label of the image may indicate that each pixel in the image belongs to one of the predetermined instance categories. For example, a semantic segmentation network may classify pixels corresponding to different persons in an image into the same semantic category, for example, a person. However, an instance segmentation network may classify these pixels into different instance categories corresponding to different persons. Hereinafter, some embodiments will be described with reference to semantic segmentation. However, it should be understood that these embodiments are also applicable to instance segmentation.

[0024] Figures 2A to 2C are schematic diagrams of image processing according to some embodiments of the present disclosure. As shown in Figures 2A to 2C, the image processing apparatus 120 may include a registration module 121 and a segmentation module 122. The segmentation module 122 may be a semantic segmentation module or an instance segmentation module. In the example shown in Figures 2A to 2C, for example, the segmentation module 122 is a semantic segmentation module. The images processed by the image processing apparatus 120 may include images 201, 202, and 203 related to nasal lesions captured under different light sources.

[0025] In FIG. 2, for example, image 201 may be selected as a target image on which semantic segmentation is performed. Images 202 and 203 are related images. The registration module 121 may register the related images 202 and 203 to the target image 201 to obtain registration images 204 and 205. In some embodiments, for each of the related images 202 and 203, the registration module 121 may generate a transformed image by registering the related image to the target image 201 based on a first transformation, and then generate a registration image by registering the transformed image to the target image 201 based on a second transformation. In some embodiments, the first transformation may be an affine transformation or a rigid transformation. Examples of the first transformation may include, but are not limited to, translation, rotation, scaling, etc. In some embodiments, the second transformation may be a deformable transformation or a non-rigid transformation. For example, the second transformation may refer to the transformation of pixels or content of an image. In some embodiments, the registration module 121 may use a trained image registration network to register the related images 202 and 203 to the target image 201 as described below with reference to FIG. 4.

[0026] As shown in FIG. 2A, the target image 201 and the registration images 204 and 205 may be input into the segmentation module 122. The segmentation module 122 may perform semantic segmentation on the images 201, 204, and 205 to generate their semantic segmentation results. The semantic segmentation may be performed by using a trained semantic segmentation network or by using any other suitable algorithm known currently or developed in the future. The segmentation module 122 may generate a final segmentation result 231 by combining the semantic segmentation results of the images 201, 204, and 205 based on the weights of the images 201, 204, and 205. In some embodiments, the respective weights associated with the images 201, 204, and 205 may be predetermined. For example, since no conversion is performed on the target image 201, the weight associated with the target image 201 may be the highest. The weight associated with the registration image 204 and the weight associated with the registration image 205 may be the same as or different from each other. The segmentation module 122 may determine the weighted sum of the semantic segmentation results of the images 201, 204, and 205 as the final segmentation result 211. The final segmentation result 211 may be input into the argmax function 123 to generate a semantic segmentation label 212 for the target image 201. For example, as described above, the semantic segmentation label 212 may indicate the respective semantic categories of the pixels within the target image 201. In some embodiments, the semantic segmentation label 212 may be provided to a doctor or an automatic diagnosis system for disease diagnosis.

[0027] In some embodiments, for the same group of images of the same object captured under different light sources, since each image can be selected as a target image, a group of segmentation labels can be generated for the group of images. In some embodiments, the group of segmentation labels may be directly provided to a doctor or an automatic diagnosis system for disease diagnosis. Alternatively, the group of segmentation labels may be combined in an appropriate manner and then provided to a doctor or an automatic diagnosis system for disease diagnosis.

[0028] As shown in FIG. 2B, for example, image 203 may be selected as the target image, and images 202 and 201 may be selected as related images. The registration module 121 may register the related images 202 and 201 to the target image 203 to obtain registration images 206 and 207. The segmentation module 122 may perform semantic segmentation on images 203, 206, and 207 and combine their semantic segmentation results into the final segmentation result 221. The final segmentation result 221 may be input to the argmax function 123 to generate the semantic segmentation label 222 for the target image 203. As shown in FIG. 2C, for example, image 202 may be selected as the target image, and images 203 and 201 may be selected as related images. The registration module 121 may register the related images 203 and 201 to the target image 202 to obtain registration images 208 and 209. The segmentation module 122 may perform semantic segmentation on images 202, 208, and 209 and combine their semantic segmentation results into the final segmentation result 231. The final segmentation result 231 may be input to the argmax function 123 to generate the semantic segmentation label 232 for the target image 203.

[0029] For example, the semantic segmentation labels 212, 222, and 232 may be provided directly to a doctor or an automatic diagnosis system for disease diagnosis. Alternatively, the semantic segmentation labels 212, 222, and 232 may be combined in a suitable manner and then provided to a doctor or an automatic diagnosis system for disease diagnosis.

[0030] In some embodiments, the image processing device 120 shown in FIG. 1 may perform an object identification task. For example, in response to acquiring a plurality of images regarding an object captured under different light sources, the image processing device 120 may register the plurality of images and identify the object based on the registered plurality of images.

[0031] FIG. 3 is a schematic diagram of image processing according to some embodiments of the present disclosure. As shown in FIG. 3, the image processing device 120 may include the same registration module 121 as in FIGS. 2A - 2C and an object identification module 124. The images processed by the image processing device 120 may include images 301, 302, and 303 regarding a nasal lesion captured under different light sources.

[0032] In FIG. 3, for example, image 301 may be selected as the target image. Images 302 and 303 are related images. The registration module 121 may register the related images 302 and 303 to the target image 301 to obtain registration images 304 and 305. Images 301, 304, and 305 may be input to the object identification module 124. In some embodiments, the object identification module 124 may identify an object based on images 301, 304, and 305. For example, the object identification module 124 may perform object identification for each of images 301, 304, and 305 by using a trained object identification network or any suitable algorithm known currently or developed in the future to obtain their respective object identification results. Then, the object identification module 124 may combine the object identification results into a final object identification result 306 based on their respective weights. Alternatively, in some embodiments, the object identification module 124 may perform semantic segmentation or instance segmentation for each of images 301, 304, and 305 to obtain their respective segmentation results, and combine the segmentation results into a final segmentation result based on their respective weights. Then, the object identification module 124 may identify an object based on the final segmentation result to obtain a final object identification result 306. For example, the final object identification result 306 may be provided to a doctor or an automatic diagnosis system for disease diagnosis.

[0033] In some embodiments, for the same group of images of the same object captured under different light sources, each image can be selected as a target image, so that a group of object identification results can be generated for the group of images. In some embodiments, the group of object identification results may be directly provided to a doctor or an automatic diagnosis system for disease diagnosis. Alternatively, the group of object identification results may be combined in an appropriate manner and then provided to a doctor or an automatic diagnosis system for disease diagnosis.

[0034] FIG. 4 is a schematic diagram of image registration according to some embodiments of the present disclosure. FIG. 4 shows an image registration network 400 that can be used in the registration network 121 shown in FIGS. 2A-2C and FIG. 3. The image registration network 400 includes subnetworks 410 and 420. The subnetwork 410 may be trained to register an image based on a first transformation. As described above, the first transformation may be an affine transformation or a rigid body transformation. Examples of the first transformation may include, but are not limited to, translation, rotation, scaling, etc. The subnetwork 410 may be trained to register an image based on a second transformation. As described above, the second transformation may be a deformable transformation or a non-rigid body transformation. For example, the second transformation may refer to the transformation of the pixels or content of an image. The image registration network 400 may be trained based on a group of training image pairs, where each training image pair includes a fixed image and a moving image relative to the fixed image. The fixed image and the moving image within the group of training image pairs may have the same number of pixels.

[0035] As shown in FIG. 4, in the training stage, each training image pair including a fixed image 402 (hereinafter also referred to as "F") and a moving image 401 (hereinafter also referred to as "M") is input into the subnetwork 410 to obtain a first transformation function

Number

Number

Number

Number

Number

Number

Number

[0036] In some embodiments, the target loss for training the image registration network 400 may be determined based on the fixed image 402, the first transformed image 403, and the second transformed image 404. The network parameters of the image registration network 400 may be iteratively updated such that the target loss is minimized. In some embodiments, the target loss for training the image registration network 400 may be determined as a weighted sum of a group of losses.

[0037] As shown in FIG. 4, in some embodiments, the first similarity loss 441 may be determined based on the fixed image 402 (i.e., F) and the first transformed image 403 (i.e.,

Number

Number

[0038] As shown in FIG. 4, in some embodiments, the second similarity loss 442 may be determined based on the fixed image 402 (i.e., F) and the second transformed image 404 (i.e.,

Number

Number

[0039] As shown in FIG. 4, in some embodiments, the third similarity loss 443 may be based on the first transformed image 403 (i.e., [Number] ) and the second transformed image 404 (i.e., [Number] ) may be determined based on. For example, the third similarity loss 443 may be expressed as the following equation. [Number] Here [Number] is. The third similarity loss [Number] is the backward similarity loss for [Number] and [Number] to improve the registration accuracy.

[0040] As shown in FIG. 4, in some embodiments, the spatial smoothness loss 444 is based on the first transformed image 403 (i.e., [Number] ) and the second transformation 432 (i.e., [Number] ) may be determined based on. For example, the spatial smoothness loss 444 may be expressed as the following equation. [Number] Here, the spatial smoothness loss is to enforce a spatially smooth deformation, [Number] A regularization constraint is provided for [it]. In some embodiments, the target loss L can be determined as the weighted sum of all the above losses, that is, as the following formula. [Number] Here, [Number] where

[0041] In view of the above, it can be seen that the embodiments of the present disclosure provide a solution for image processing. According to some embodiments of the present disclosure, by synthesizing the segmentation results of images regarding an object captured under different light sources to obtain the final segmentation result for the target image, the accuracy of image segmentation (e.g., semantic segmentation or instance segmentation) for the target image can be improved. Additionally, by removing the influence of slight deformations of the object between different images through image registration, the accuracy of image segmentation and / or object identification can be improved.

[0042] FIG. 5 is a diagram showing an exemplary method 500 for image processing according to some embodiments of the present disclosure. The method 500 can be implemented in the image processing apparatus 120 as shown in FIG. 1. The method 500 may include additional blocks not shown and / or may omit some of the blocks shown, and it should be understood that the scope of the present disclosure is not limited in this regard.

[0043] In block 510, the image processing apparatus 120 may acquire a plurality of images regarding the same object captured under different light sources, where the plurality of images includes a target image and at least one related image.

[0044] In block 520, the image processing apparatus 120 may generate a segmentation label for the target image based on the segmentation results of a plurality of images.

[0045] In some embodiments, method 500 further includes generating at least one registration image by registering the at least one related image to the target image, and generating the segmentation result by performing semantic segmentation or instance segmentation on the target image and the at least one registration image.

[0046] In some embodiments, generating the at least one registration image includes generating a transformed image by registering each related image of the at least one related image to the target image based on a first transformation, and generating a registration image by registering the transformed image to the target image based on a second transformation.

[0047] In some embodiments, the first transformation is an affine transformation and the second transformation is a deformable transformation.

[0048] In some embodiments, registering the at least one related image to the target image includes registering the at least one related image to the target image by using a trained image registration network.

[0049] In some embodiments, method 500 may further include training the image registration network based on a group of training image pairs, where each training image pair includes a fixed image and a moving image with respect to the fixed image.

[0050] In some embodiments, training the image registration network includes generating a first transformed image by registering the moving image to the fixed image based on an affine transformation, generating a second transformed image by registering the first transformed image to the fixed image based on a deformable transformation, determining a target loss for training the image registration network based on the fixed image, the first transformed image, and the second transformed image, and training the image registration network such that the target loss is minimized.

[0051] In some embodiments, determining the target loss includes determining a first similarity loss based on the fixed image and the first transformed image, determining a second similarity loss based on the fixed image and the second transformed image, determining a spatial smoothness loss based on the first transformed image and a function corresponding to the deformable transformation, determining a third similarity loss based on the first transformed image and the second transformed image, and determining the target loss based on a weighted sum of the first similarity loss, the second similarity loss, the spatial smoothness loss, and the third similarity loss.

[0052] In some embodiments, performing semantic segmentation or instance segmentation on the target image and the at least one registered image includes performing the semantic segmentation on the target image and the at least one registered image by using a trained semantic segmentation network, or performing the instance segmentation on the target image and the at least one registered image by using a trained instance segmentation network.

[0053] In some embodiments, generating the semantic segmentation label for the target image includes determining a final segmentation result for the target image based on a weighted sum of the segmentation results, and generating the segmentation label for the target image based on the final segmentation result.

[0054] In some embodiments, different light sources are associated with different wavelengths or different combinations of wavelengths.

[0055] FIG. 6 is a diagram illustrating an exemplary method 600 of image processing according to some embodiments of the present disclosure. The method 600 can be implemented in the image processing apparatus 120 as shown in FIG. 1. The method 500 may include additional blocks not shown and / or may omit some of the blocks shown, and it should be understood that the scope of the present disclosure is not limited in this regard.

[0056] In block 610, the image processing apparatus 120 may acquire a plurality of images of the same object captured under different light sources.

[0057] In block 620, the image processing apparatus 120 may register the plurality of images.

[0058] In block 630, the image processing apparatus 120 may identify an object based on the registered plurality of images.

[0059] In some embodiments, the plurality of images includes a target image and at least one related image, and registering the plurality of images includes generating at least one registered image by registering the at least one related image to the target image, where the registered plurality of images includes the at least one registered image and the target image.

[0060] In some embodiments, identifying the object based on the plurality of registered images includes generating a segmentation result by performing semantic segmentation or instance segmentation on the target image and the at least one registered image, and identifying the object based on the segmentation result.

[0061] In some embodiments, generating the at least one registered image includes generating a transformed image by registering the at least one associated image to the target image based on a first transformation for each associated image of the at least one associated image, and generating a registered image by registering the transformed image to the target image based on a second transformation.

[0062] In some embodiments, the first transformation is an affine transformation and the second transformation is a deformable transformation.

[0063] In some embodiments, registering the at least one associated image to the target image includes registering the at least one associated image to the target image by using a trained image registration network.

[0064] In some embodiments, method 600 may further include training the image registration network based on a set of training image pairs, where each training image pair includes a fixed image and a moving image relative to the fixed image.

[0065] In some embodiments, training the image registration network includes generating a first transformed image by registering the moving image to the fixed image based on an affine transformation, generating a second transformed image by registering the first transformed image to the fixed image based on a deformable transformation, determining a target loss for training the image registration network based on the fixed image, the first transformed image, and the second transformed image, and training the image registration network such that the target loss is minimized.

[0066] In some embodiments, determining the target loss includes determining a first similarity loss based on the fixed image and the first transformed image, determining a second similarity loss based on the fixed image and the second transformed image, determining a spatial smoothness loss based on the first transformed image and a function corresponding to the deformable transformation, determining a third similarity loss based on the first transformed image and the second transformed image, and determining the target loss based on a weighted sum of the first similarity loss, the second similarity loss, the spatial smoothness loss, and the third similarity loss.

[0067] In some embodiments, performing semantic segmentation or instance segmentation on the target image and the at least one registered image includes performing the semantic segmentation on the target image and the at least one registered image by using a trained semantic segmentation network, or performing the instance segmentation on the target image and the at least one registered image by using a trained instance segmentation network.

[0068] In some embodiments, generating the semantic segmentation label for the target image includes determining a final segmentation result for the target image based on a weighted sum of the segmentation results, and generating the segmentation label for the target image based on the final segmentation result.

[0069] In some embodiments, different light sources are associated with different wavelengths or different combinations of wavelengths.

[0070] FIG. 7 is a schematic block diagram of an apparatus 700 that can be used to implement embodiments of the present disclosure. For example, the image collection device 110 and / or the image processing device 120 can be implemented by the apparatus 700. For example, the apparatus 700 can be used to implement a medical assistance device or an endoscope assistance device that can capture images of suspected lesions or tumors under different light sources. As shown in FIG. 7, the apparatus 700 includes a central processing unit (CPU) 701 that can execute various appropriate operations and processes based on computer program instructions stored in a read-only memory (ROM) 702 or computer program instructions uploaded from a storage unit 708 to a random access memory (RAM) 703. The RAM 703 further stores various programs and data required for the operation of the apparatus 700. The CPU 701, the ROM 702, and the RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0071] The I / O interface 705 is connected to components including an input unit 706 such as a keyboard and a mouse, an output unit 707 such as various displays and speakers, a storage unit 708 such as a magnetic disk and an optical disk, and a communication unit 709 such as a network card, a modem, and a wireless communication transceiver. The communication unit 709 enables the device 700 to exchange data / information with other devices via a computer network such as the Internet and / or a telecommunications network.

[0072] The above-described method or process, for example, methods 500 and / or 600, can be executed by the processing unit 701. For example, in some implementations, method 500 may be implemented as a computer software program tangibly included in a machine-readable medium such as the storage unit 708. In some implementations, the computer program may be partially or fully loaded and / or mounted on the device 700 by the ROM 702 and / or the communication unit 709. When the computer program is uploaded to the RAM 703 and executed by the CPU 701, one or more steps of the above-described methods 500 and / or 600 can be executed.

[0073] In some embodiments, the image processing apparatus includes a circuit, and the circuit is configured to obtain a plurality of images regarding an object captured under different light sources, the plurality of images including a target image and at least one related image, and generate a segmentation label for the target image based on the segmentation results of the plurality of images.

[0074] In some embodiments, the circuit is further configured to generate at least one registered image by registering the at least one related image to the target image, and generate the segmentation result by performing semantic segmentation or instance segmentation on the target image and the at least one registered image.

[0075] In some embodiments, the image processing apparatus includes a circuit configured to acquire a plurality of images of an object captured under different light sources, register the plurality of images, and identify the object based on the registered plurality of images.

[0076] In some embodiments, the plurality of images includes a target image and at least one related image, and the circuit is further configured to generate at least one registered image by registering the at least one related image to the target image, where the registered plurality of images includes the at least one registered image and the target image.

[0077] In some embodiments, the circuit is further configured to generate a segmentation result by performing semantic segmentation or instance segmentation on the target image and the at least one registered image, and identify the object based on the segmentation result.

[0078] In some embodiments, the circuit is further configured to generate a transformed image by registering each of the at least one associated image to the target image based on a first transformation, and to generate a registered image by registering the transformed image to the target image based on a second transformation.

[0079] In some embodiments, the first transformation is an affine transformation and the second transformation is a deformable transformation.

[0080] In some embodiments, the circuit is further configured to register the at least one associated image to the target image by using a trained image registration network.

[0081] In some embodiments, the circuit is further configured to train the image registration network based on a group of training image pairs, where each training image pair includes a fixed image and a moving image relative to the fixed image.

[0082] In some embodiments, the circuit is further configured to generate a first transformed image by registering the moving image to the fixed image based on an affine transformation, to generate a second transformed image by registering the first transformed image to the fixed image based on a deformable transformation, to determine a target loss for training the image registration network based on the fixed image, the first transformed image, and the second transformed image, and to train the image registration network such that the target loss is minimized.

[0083] In some embodiments, the circuit is further configured to determine a first similarity loss based on the fixed image and the first transformed image, determine a second similarity loss based on the fixed image and the second transformed image, determine a spatial smoothness loss based on the first transformed image and a function corresponding to the deformable transformation, determine a third similarity loss based on the first transformed image and the second transformed image, and determine the target loss based on a weighted sum of the first similarity loss, the second similarity loss, the spatial smoothness loss, and the third similarity loss.

[0084] In some embodiments, the circuit is further configured to perform the semantic segmentation on the target image and the at least one registration image by using a trained semantic segmentation network, or perform the instance segmentation on the target image and the at least one registration image by using a trained instance segmentation network.

[0085] In some embodiments, the circuit is further configured to determine a final segmentation result for the target image based on a weighted sum of the segmentation results, and generate the segmentation label for the target image based on the final segmentation result.

[0086] In some embodiments, different light sources are associated with different wavelengths or different combinations of wavelengths.

[0087] The present disclosure may be implemented as a system, method, and / or computer program product. When the present disclosure is implemented as a system, in addition to being realized on a single device, the components described herein may also be realized in the form of a cloud computing architecture. In a cloud computing environment, these components are remotely located and can operate together to realize the functions described in the present disclosure. Cloud computing can provide computing, software, data access, and storage services. An end user need not know the physical location or settings of the systems or hardware that provide these services. Cloud computing can provide services over a wide area network (e.g., the Internet) using appropriate protocols. For example, a cloud computing provider can provide applications over a wide area network, and these can be accessed via a browser or any other computing component. Cloud computing components and corresponding data can be stored on a remote server. The computing resources within a cloud computing environment may be centralized in a remote data center or these computing resources may be distributed. A cloud computing infrastructure can provide services via these shared data centers even though the shared data centers may appear as a single access point to a user. Thus, a cloud computing architecture can be used to provide the various functions described herein from a remote service provider. Alternatively, these functions may be provided from a conventional server and may be installed directly or otherwise on a client device. Additionally, the present disclosure may be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for executing various aspects of the present disclosure.

[0088] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, punch card, or mechanically encoded device such as a raised structure in a groove in which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as being, in itself, a transient signal such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical cable), or an electrical signal transmitted via a wire.

[0089] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or to an external computer or external storage device via a network such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in the computer-readable storage medium within each respective computing / processing device.

[0090] Computer-readable program instructions for executing the operations of the present disclosure may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for implementing aspects of the present disclosure.

[0091] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying flowchart diagrams and / or block diagrams. It should be understood that each block of the flowchart diagrams and / or block diagrams, and combinations of blocks in the flowchart diagrams and / or block diagrams, can be implemented by computer-readable program instructions.

[0092] These computer-readable program instructions are provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate means for realizing the functions / operations specified within one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium storing the instructions comprises a product including instructions for realizing the aspects of the functions / operations specified within one or more blocks of a flowchart and / or block diagram.

[0093] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are executed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process whereby the instructions executed on the computer, other programmable data processing apparatus, or other device realize the functions / operations specified in the flowchart and / or block diagram blocks.

[0094] Flowcharts and block diagrams show the architecture, functionality, and operation of possible realizations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, snippet, or portion of code that includes one or more executable instructions for implementing a specified logical function. In some alternative realizations, the functions recorded in the blocks may occur in an order different from the order recorded in the figures. For example, depending on the related functions, two blocks shown in succession may actually be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order. It should also be noted that each block in the block diagrams and / or flowchart diagrams, as well as combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs a particular function or operation, or by a combination of dedicated hardware and computer instructions.

[0095] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limiting of the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application to technologies found in the marketplace, or the technical improvement thereof, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for image processing, comprising: obtaining a plurality of images related to an object captured under different light sources, the plurality of images including a target image and at least one related image; generating a segmentation label for the target image based on the segmentation results of the plurality of images; generating at least one registered image by registering the at least one related image to the target image; generating the segmentation result by performing semantic segmentation or instance segmentation on the target image and the at least one registered image; wherein registering the at least one related image to the target image includes registering the at least one related image to the target image by using a trained image registration network; the method further includes training the image registration network based on a group of training image pairs, each training image pair including a fixed image and a moving image relative to the fixed image; training the image registration network includes generating a first transformed image by registering the moving image to the fixed image based on an affine transformation; generating a second transformed image by registering the first transformed image to the fixed image based on a deformable transformation; determining a target loss for training the image registration network based on the fixed image, the first transformed image, and the second transformed image; training the image registration network such that the target loss is minimized; wherein determining the target loss includes determining a first similarity loss based on the fixed image and the first transformed image; determining a second similarity loss based on the fixed image and the second transformed image; determining a spatial smoothness loss based on the first transformed image and a function corresponding to the deformable transformation; determining a third similarity loss based on the first transformed image and the second transformed image; Determining the target loss based on a weighted sum of the first similarity loss, the second similarity loss, the spatial smoothness loss, and the third similarity loss; A method comprising the above. **Claim 2** An image processing method, comprising: Obtaining a plurality of images regarding an object captured under different light sources; Registering the plurality of images; Identifying the object based on the registered plurality of images; Including: The plurality of images include a target image and at least one related image, and registering the plurality of images includes: Generating at least one registered image by registering the at least one related image to the target image; The registered plurality of images include the at least one registered image and the target image; Registering the at least one related image to the target image includes: Registering the at least one related image to the target image by using a trained image registration network; The method further includes: Training the image registration network based on a group of training image pairs; Each training image pair includes a fixed image and a moving image relative to the fixed image; Training the image registration network includes: Generating a first transformed image by registering the moving image to the fixed image based on an affine transformation; Generating a second transformed image by registering the first transformed image to the fixed image based on a deformable transformation; Determining a target loss for training the image registration network based on the fixed image, the first transformed image, and the second transformed image; Training the image registration network such that the target loss is minimized; Including: Determining the target loss includes: Determining a first similarity loss based on the fixed image and the first transformed image; Determining a second similarity loss based on the fixed image and the second transformed image; Determining a spatial smoothness loss based on the first transformed image and a function corresponding to the deformable transformation; Determining a third similarity loss based on the first transformed image and the second transformed image; Determining the target loss based on a weighted sum of the first similarity loss, the second similarity loss, the spatial smoothness loss, and the third similarity loss; A method comprising the above.

3. Identifying the object based on the registered plurality of images comprises: Generating a segmentation result by performing semantic segmentation or instance segmentation on the target image and the at least one registered image; Identifying the object based on the segmentation result. The method according to claim 2, comprising the above.

4. Generating the at least one registered image comprises: For each of the at least one associated image, Generating a transformed image by registering the at least one associated image to the target image based on a first transformation; Generating a registered image by registering the transformed image to the target image based on a second transformation. The method according to claim 1 or 2, comprising the above.

5. The first transformation is an affine transformation and the second transformation is a deformable transformation. The method according to claim 4.

6. Performing semantic segmentation or instance segmentation on the target image and the at least one registered image comprises: Performing the semantic segmentation on the target image and the at least one registered image by using a trained semantic segmentation network, or Performing the instance segmentation on the target image and the at least one registered image by using a trained instance segmentation network. The method according to claim 1 or 3, comprising the above.

7. Generating the segmentation label for the target image comprises: Determining a final segmentation result for the target image based on the weighted sum of the segmentation results; Generating the segmentation label for the target image based on the final segmentation result; The method according to claim 1, comprising the above.

8. The different light sources are associated with different wavelengths or different combinations of wavelengths The method according to any one of claims 1 to 7.

9. An image processing apparatus comprising at least one processor, The at least one processor Obtaining a plurality of images regarding an object captured under different light sources, the plurality of images including a target image and at least one associated image; Generating a segmentation label for the target image based on the segmentation results of the plurality of images; Generating at least one registration image by registering the at least one associated image to the target image; Generating the segmentation result by performing semantic segmentation or instance segmentation on the target image and the at least one registration image; Registering the at least one associated image to the target image by using a trained image registration network; Configured to train the image registration network based on a group of training image pairs, Each training image pair includes a fixed image and a moving image with respect to the fixed image, The at least one processor further Generating a first transformed image by registering the moving image to the fixed image based on an affine transformation; Generating a second transformed image by registering the first transformed image to the fixed image based on a deformable transformation; Determining a target loss for training the image registration network based on the fixed image, the first transformed image, and the second transformed image; Training the image registration network such that the target loss is minimized; Determining a first similarity loss based on the fixed image and the first transformed image; Determine a second similarity loss based on the fixed image and the second transformed image, Determine a spatial smoothness loss based on the first transformed image and the function corresponding to the deformable transformation, Determine a third similarity loss based on the first transformed image and the second transformed image, Determine the target loss based on a weighted sum of the first similarity loss, the second similarity loss, the spatial smoothness loss, and the third similarity loss An image processing apparatus configured as described above.

10. An image processing apparatus including at least one processor, wherein the at least one processor, acquires a plurality of images regarding an object captured under different light sources, registers the plurality of images, is configured to identify the object based on the registered plurality of images, the plurality of images includes a target image and at least one associated image, the at least one processor further, is configured to generate at least one registered image by registering the at least one associated image to the target image, the registered plurality of images includes the at least one registered image and the target image, the at least one processor further, registers the at least one associated image to the target image by using a trained image registration network, is configured to train the image registration network based on one group of training image pairs, each training image pair includes a fixed image and a moving image with respect to the fixed image, the at least one processor further, generates a first transformed image by registering the moving image to the fixed image based on an affine transformation, generates a second transformed image by registering the first transformed image to the fixed image based on a deformable transformation, determines a target loss for training the image registration network based on the fixed image, the first transformed image, and the second transformed image, trains the image registration network so that the target loss is minimized, determines a first similarity loss based on the fixed image and the first transformed image, Determine a second similarity loss based on the fixed image and the second transformed image, Determine a spatial smoothness loss based on the first transformed image and the function corresponding to the deformable transformation, Determine a third similarity loss based on the first transformed image and the second transformed image, Determine the target loss based on a weighted sum of the first similarity loss, the second similarity loss, the spatial smoothness loss, and the third similarity loss An image processing apparatus configured as described above. **Claim 11** The at least one processor further Performs semantic segmentation or instance segmentation on the target image and the at least one registration image to generate a segmentation result, Identifies the object based on the segmentation result The apparatus according to claim 10, configured as described above. **Claim 12** The at least one processor further For each of the at least one associated image among the at least one associated images, Generates a transformed image by registering the at least one associated image to the target image based on a first transformation, Generates a registration image by registering the transformed image to the target image based on a second transformation The apparatus according to claim 9 or 10, configured as described above. **Claim 13** The first transformation is an affine transformation, and the second transformation is a deformable transformation The apparatus according to claim 12. **Claim 14** The at least one processor further Performs the semantic segmentation on the target image and the at least one registration image by using a trained semantic segmentation network, or Performs the instance segmentation on the target image and the at least one registration image by using a trained instance segmentation network The apparatus according to claim 9 or 11, configured as described above. **Claim 15** The at least one processor further Determines a final segmentation result for the target image based on a weighted sum of the segmentation results Generating the segmentation label for the target image based on the final segmentation result The apparatus according to claim 9, which is set to be [

16. ] The different light sources are associated with different wavelengths or different combinations of wavelengths The apparatus according to any one of claims 9 to 15

Citation Information

Patent Citations

  • Full-automatic registration and segmentation method for multi-parameter magnetic resonance image

    CN111260700A

  • Acquisition of surface region arrangement of object by region division of image

    JP2006030106A

  • Endoscope system, image processing device, operation method of image processing device

    JP2017192565A

  • Computer scoring based on primary staining and immunohistochemistry images

    JP2020502534A