Alignment of images acquired using different endoscope units

JP2024533111A5Pending Publication Date: 2025-09-02SURGVISION GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024513706
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-16
Filing Date
2022-09-12
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Correlating information from reflectance and fluorescence images acquired by different endoscopic units is difficult due to differences in shape, zooming, resolution, contrast, and noise, and misalignment caused by patient movements, leading to erroneous medical decisions.

Method used

A method for registering images using different endoscopic units by aligning image sequences based on independently estimated probe movements, employing a neural network for automatic registration, and combining images on a monitor for accurate representation.

Benefits of technology

Facilitates accurate alignment and combination of fluorescence and reflectance images, improving medical procedure outcomes and patient health by providing a precise and aligned image representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A technique is proposed for imaging a body part (103) of a patient (106) using an endoscope system (100). A corresponding method (500) comprises steps (509-515) of acquiring two series of images using different probes (127n, 127b) of corresponding endoscope units (115m, 115b) movable between them. Corresponding images of each pair in the two sequences are registered (518-584) according to corresponding movements of these probes (127n, 127b), which are estimated, for example, independently according to the corresponding images. A method (800) is also proposed for training a neural network (439) that can be used to register the images. Corresponding computer programs (400; 700) and computer program products for operating the endoscope system (100) and for training the neural network (439) are proposed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to the field of medical devices. More specifically, the present disclosure relates to endoscopes. [Background technology]

[0002] The background of the present disclosure is presented below along with a discussion of technologies relevant to its context. However, even if this discussion refers to documents, acts, works, or the like, this does not imply or represent that the technologies discussed are part of the prior art or are common general knowledge in the field relevant to the present disclosure.

[0003] Endoscopes are commonly used in several medical procedures (e.g., endoscopy) to image parts of a patient's body that are normally invisible, for example for diagnostic, surgical and therapeutic purposes (e.g., gastroscopy to investigate conditions affecting the digestive tract and possibly remove tumors).

[0004] For this purpose, an endoscope comprises an elongated probe (generally flexible) that is inserted (via a natural orifice or a small incision) directly into a cavity defined by a body part of a patient. The tip of the probe is used to reach an area of ​​interest of the body part (in the example in question, a part of the digestive tract with a possible tumor) for a corresponding medical procedure. The tip of the probe makes it possible to illuminate the area of ​​interest and to obtain an image thereof, which is provided to a monitor for its display. The probe also has one or more working channels, through which various tools can be inserted for use during the medical procedure (for example to take a sample of a (suspected) tumor or to resect a (identified) tumor). The endoscope allows the physician to observe the inside of the cavity during the medical procedure, while at the same time with minimal trauma for the patient.

[0005] Recently, endoscopic techniques based on fluorescence imaging have been introduced into clinical studies (e.g. to improve early detection of tumors). Fluorescence imaging exploits the phenomenon of fluorescence, which occurs in fluorescent substances (called fluorophores) that emit fluorescent light when illuminated. In this case, a fluorescent agent (configured to reach and become fixed to a specific target (in this case the tumor) is administered to the patient. Images of the region of interest in the body part are then acquired and provided to a monitor for display. In this case, the images are defined by the fluorescent light emitted from the region of interest in the body part and represent the fluorophores present there. Thus, the representation of the (immobilized) fluorescent agent in these images facilitates the identification (and quantification) of the target.

[0006] It is also possible to use the above mentioned endoscopic techniques, respectively distinguished as standard (endoscopic) and fluorescent (endoscopic) techniques, together: in particular, a standard technology based endoscope can be used to acquire (reflectance) images and display them on a monitor, and a fluorescent technology based endoscope can be used to acquire (fluorescence) images and display them on a monitor.

[0007] Furthermore, a simplified endoscope unit based on fluorescence technology (called a baby scope) can be inserted into the working channel of an endoscope based on standard technology (called a mother scope). This is cost-effective, since it allows the use of an already available mother scope with the addition of a (simpler) baby scope that utilizes other features of the mother scope (e.g. steering in the cavity). Furthermore, this allows the combination of the chip-on-chip imaging technology used in the mother scope (videoscope) with the imaging via fiber bundles used in the baby scope (providing improved sensitivity). Furthermore, the same mother scope (already inserted in the patient's cavity) can be used for other different medical procedures. For example, it is possible to start inserting the baby scope into the working channel of the mother scope to analyze a possible tumor. Afterwards, the baby scope is removed and another tool is inserted into the same working channel, or the baby scope is kept in the mother scope and another tool is inserted into an additional working channel. Said further tool is used, for example, to take a sample of the (suspected) tumor or to resect the (identified) tumor. In this case as well, the reflectance image acquired by the Mother Scope is displayed on the monitor, and the fluorescence image acquired by the Baby Scope is displayed on the monitor.

[0008] However, correlation of information provided by reflectance and fluorescence images displayed on different monitors can be difficult. This is further hindered by the fact that reflectance and fluorescence images differ significantly in terms of, for example, shape, zooming, resolution, contrast, noise, etc. Furthermore, since the baby scope is generally not fixed in the working channel of the mother scope, unknown movements (e.g., back-and-forth rotation, translation) between the baby scope and the mother scope occur during the medical procedure, for example due to peristalsis and unavoidable movements of the patient. As a result, corresponding misalignments occur between the luminescence and reflectance images, leading to misinterpretation of the physician. Summary of the Invention [Means for solving the problem]

[0009] A simplified summary of the disclosure is presented here in order to provide a basic understanding of the disclosure. However, its sole purpose is to introduce some concepts of the disclosure in a simplified form as a prelude to the more detailed description below. It is not intended to identify key elements or to delineate the scope of the disclosure.

[0010] In general terms, the present disclosure is based on the idea of ​​registering images acquired with different endoscopic units.

[0011] In particular, one aspect provides a method for imaging a body part of a patient using an endoscope system, the method comprising acquiring two image sequences using different probes (of corresponding endoscope units) that are movable between them, and corresponding images of each pair in the two sequences are registered according to corresponding movements of the probes, which are, for example, estimated independently according to the corresponding images.

[0012] A further aspect provides a method for training a neural network that can be used to align images.

[0013] A further aspect provides a computer program product for operating an endoscopic system.

[0014] A further aspect provides a corresponding computer program product.

[0015] A further aspect provides a computer program for training a neural network.

[0016] A further aspect provides a corresponding computer program product.

[0017] A further aspect provides a corresponding endoscopic system for imaging a body part.

[0018] A further aspect provides an endoscopic device including one of the endoscopic units.

[0019] A further aspect provides a computing device for operating an endoscopic system.

[0020] A further aspect provides a computer system for training a neural network.

[0021] A further aspect provides a corresponding surgical method.

[0022] A further aspect provides a corresponding diagnostic method.

[0023] Further aspects provide corresponding methods of treatment.

[0024] More particularly, one or more aspects of the present disclosure are set forth in independent claims and advantageous features thereof are set forth in dependent claims, the language of all claims being incorporated herein by reference in their entirety (with any advantageous feature provided with reference to any particular aspect applying mutatis mutandis to all other aspects). [Brief description of the drawings]

[0025] The techniques of the present disclosure, and additional features and advantages thereof, will be best understood by reference to the following detailed description, given solely by way of a non-limiting indication when read in conjunction with the accompanying drawings, in which, for simplicity, corresponding elements are given equal or similar reference numerals, their descriptions will not be repeated, and the name of each entity will generally be used to indicate both its type and its attributes (e.g., value, content, and expression) in which:

[0026] [Figure 1] 1 shows a diagrammatic representation of an endoscopic system according to one embodiment of the present disclosure. [Diagram 2]FIG. 1 shows a functional block diagram of an endoscopic system that can be used to implement a technique according to an embodiment of the present disclosure. [Diagram 3] 1 illustrates the general principle of the approach according to one embodiment of the present disclosure. [Figure 4] 1 illustrates major software components that can be used to implement a technique according to one embodiment of the present disclosure. [Diagram 5] 1 shows an operational diagram illustrating the flow of operations associated with implementing a technique according to one embodiment of the present disclosure. [Figure 6] FIG. 1 shows a schematic block diagram of a computer system that can be used to train a neural network of a method according to an embodiment of the present disclosure. [Figure 7] 1 illustrates the main software components that can be used to train a neural network in a technique according to one embodiment of the present disclosure. [Figure 8] FIG. 1 shows an operational diagram illustrating the flow of operations associated with training a neural network in a technique according to one embodiment of the present disclosure. [Figure 9A] 1 illustrates various application examples of a technique according to an embodiment of the present disclosure. [Figure 9B] 1 illustrates various application examples of a technique according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0027] With particular reference to FIG. 1, a diagrammatic representation of an endoscopic system 100 is shown in accordance with one embodiment of the present disclosure.

[0028] The endoscopic system, or simply endoscope 100, is used in medical procedures to image a body part 103 of a patient 106 within a cavity defined thereby (the cavity of the body part 103 is not normally visible). The body part 103 includes a target of the medical procedure, e.g. a lesion such as a tumor 109, to be examined, removed or treated. The cavity of the body part 103 is accessible via an opening 112, which may be a natural orifice of the patient 106 or a small incision in the skin. For example, in diagnostic applications, the endoscope 100 allows to find / monitor a lesion, in (minimally invasive) surgical applications, the endoscope 100 allows to identify a lesion to be removed, and in therapeutic applications, the endoscope 100 allows to depict a lesion to be treated. Examples of these medical procedures are gastroscopy, colonoscopy, esophagoscopy, etc. for diagnostic applications, arthroscopy, laparoscopy, thoracoscopy, etc. for surgical applications, and cauterization, dilation, stent placement, etc. for therapeutic applications.

[0029] The endoscope 100 has an unconventional structure that is used to apply fluorescent (endoscopic) techniques (e.g. to display fluorescent substances, such as fluorescent agents adapted to accumulate in tumors, pre-administered to the patient 106) in combination with standard (endoscopic) techniques (to display what is visible to the human eye). For this purpose, the endoscope 100 is composed of two endoscopic units, namely a (first) main endoscopic unit 115m (which itself is a conventional device used in clinical practice) and a (second) auxiliary endoscopic unit 115b.

[0030] The main endoscope unit, or simply the mother scope 115m, comprises the following components: A central unit of the mother scope 115m is used to manage its operation. For example, the central unit is implemented as a trolley 118m, with four castors arranged at its corresponding lower corners to facilitate its movement (foot brakes (not shown) are provided to fix the trolley 118m in place). A monitor 121m (for example mounted on the top of the trolley 118m) is used to display images of the body part 103 during the medical procedure. A video interface 124m (for example a serial digital interface (SDI) port at the back of the trolley 118m) is used to exchange video information with the outside. A probe 127m (for example connected to the trolley 118m via a cable) is used to act on the patient 106. For example, the probe 127m is implemented as an elongated shaft for insertion into a cavity of the body part 103. The shaft of the probe 127m may be rigid or preferably flexible, allowing it to slide through the cavity of the body part 103 even if the cavity of the body part 103 has a curved path. The distal end, or tip 130m, of the probe 127m is used to reach the area of ​​interest of the medical procedure in the cavity of the body part 103, to illuminate it and to acquire its (reflected) image (as described below). At the proximal end of the probe 127m (outside the cavity 103) a handle 133m is provided (via a control cable, not shown) for driving the tip 130m. The probe 127m has one or more working channels (accessible via one or more corresponding working ports close to its proximal end), only one of which is shown in the figures with reference number 136m. The working channels allow the insertion of various tools (e.g., snares, forceps, knives, clip appliers, etc.) used during the medical procedure. A dedicated working channel can be connected to a fluid injector / extractor (not shown) for irrigating the tip 130m and the cavity of the body part 103 during a medical procedure.

[0031] The auxiliary endoscope unit, or simply baby scope 115b, comprises the following components: As mentioned above, a central unit of the baby scope 115b is used to manage its operation. For example, the central unit is implemented as a dolly 118b (with four casters to facilitate its movement and a foot brake to fix the dolly 118b in place). A monitor 121b (for example mounted on top of the dolly 118b) is used to display (further) images of the cavity of the body part 103 during the medical procedure. A video interface 124b (for example an SDI port on the back of the dolly 118m) is used to exchange video information with the outside world. A probe 127b (for example connected to the dolly 118b via a cable) is used to act on the patient 106. For example, the probe is implemented as an elongated shaft (preferably flexible). The distal end or tip 130b of probe 127b is used to reach, illuminate and acquire (fluorescence and possibly reflectance) images of the same area of ​​interest for the medical procedure within the cavity of body part 103 (as described below).

[0032] In a particular implementation of the technique according to an embodiment of the present disclosure, probe 127b is thinner than probe 127m. Probe 127b is inserted into working channel 136m until its tip 130b reaches tip 130m of probe 127m (without compromising the maneuverability of probe 127m due to its size and flexibility). Furthermore, carriage 118m and carriage 118b are connected to each other, for example, via cable 139, to exchange information during the medical procedure.

[0033] Referring now to FIG. 2, a functional block diagram of an endoscope 100 that can be used to perform techniques in accordance with embodiments of the present disclosure is shown.

[0034] Starting from the baby scope 115b, it comprises the following components: An illumination unit is used to illuminate the area of ​​interest of the cavity 103. For this purpose, an excitation light source 203b (inside the carriage of the baby scope 115b, not shown), for example of the laser type, generates excitation light for the fluorescent substances. In particular, the excitation light has a wavelength and energy suitable for exciting the phosphors of the fluorescent agents (for example of the near infrared (NIR) type). Optionally, a white light source 209b (inside the carriage), for example of the xenon type, generates white light (which appears substantially colorless to humans, for example containing all wavelengths of the spectrum visible to the human eye at the same intensity). This white light is mixed together with the excitation light. A delivery optics 212b (at the tip of the probe of the baby scope 115b, not shown) delivers the excitation light and possibly the white light to the area of ​​interest of the body part 103. An incoherent optical fiber bundle 218b (along the probe) transmits the excitation light from the excitation light source 203b, possibly mixed with the white light from the white light source 209b, to the delivery optics 212b (alternatively, not shown, two separate delivery optics with corresponding incoherent optical fiber bundles are used to deliver the excitation light and the white light independently). The acquisition unit is used to acquire (fluorescence and possibly reflection) images of the region of interest of the body part 103 within its field of view 221b (i.e. the part of the world within the solid angle that the acquisition unit can sense), which in the example under consideration includes the target 109. For this purpose, the collection optics 224b collects light from the field of view 221b (epi-illumination geometry). The collected light comprises fluorescent light emitted by fluorescent molecules present within the field of view 221b (illuminated by the excitation light). Indeed, when a fluorescent molecule absorbs the excitation light, it goes into an excited (electronic) state, from which it decays to the ground (electronic) state in a very short time, as the excited state is unstable, thereby emitting fluorescence (at a characteristic wavelength longer than that of the excitation light, due to the energy dissipated as heat in the excited state), the intensity of which mainly depends on the amount of fluorescent molecule irradiated. Moreover, the collected light includes visible light (within the visible spectrum) reflected by any objects present in the field of view 221b (irradiated by white light).A beam splitter 227b splits the collected light into two channels. For example, the beam splitter 227b is a dichroic mirror that transmits and reflects respectively the collected light at wavelengths above and below a threshold wavelength between the spectrum of visible light and the spectrum of fluorescent light (or vice versa). A bundle of coherent optical fibers 230b (along the probe) transmits the collected light from the collection optics 224b to the beam splitter 227b. In the (transmission) channel of the beam splitter 227b with the fluorescent light defined by the part of the collected light in its spectrum, an emission filter 233b filters the fluorescent light to remove residual components outside the spectrum of the fluorescent light. A fluorescence camera 236b (for example of the EMCCD type) receives the fluorescent light from the emission filter 233b and generates a corresponding fluorescence (digital) image representing the distribution of fluorescent molecules in the field of view 221b. In the other (reflected) channel of beam splitter 227b, with visible light defined by its spectrum, reflectance, or the portion of the light recovered in a photograph, as appropriate, a camera 239b (e.g. of the CCD type) receives the visible light and generates a corresponding reflectance (digital) image representative of what the human eye sees within the field of view 221b.

[0035] Turning to the mother scope 115m, it comprises the following components: As mentioned above, the illumination unit is used to illuminate the area of ​​interest of the body part 103. For this purpose, a white light source 209m (inside the carriage of the mother scope 115b, not shown), for example of the xenon type, generates white light. The delivery optics 212m (at the tip of the probe of the mother scope 115m, not shown) delivers the white light to the area of ​​interest of the body part 103. A bundle of incoherent optical fibers 218m (along the probe) transmits the white light from the white light source 209m to the delivery optics 212m. Using the acquisition unit, it acquires (reflectance) images of the area of ​​interest of the body part 103 within its field of view 221m, which in the example under consideration includes the target 109. For this purpose, the recovery optics 224m recovers the visible light (within the visible spectrum) reflected by any object present within the field of view 221m illuminated by the white light (epi-illumination geometry). A reflectance or photo camera 239m (e.g., of the CCD type) receives the visible light and produces a corresponding reflectance (digital) image representative of what the human eye sees within the field of view 221m. In a videoscope configuration, the reflectance camera 239m is located at the tip of the probe. In this case, a digital connection 240m transmits the reflectance image to the dolly (alternatively, not shown, the reflectance camera could be located inside the dolly and a bundle of coherent optical fibers could transmit visible light from the collection optics to the reflectance camera).

[0036] The central units 242b and 242m are used to control the operation of the baby scope 115b and the mother scope 115m, respectively. The central units 242b, 242m comprise several units connected between them via bus structures 245b, 245m. In particular, microprocessors (μP) 248b, 248m and others provide the logic functions of the central units 242b, 242m. The non-volatile memories (ROM) 251b, 251m store the basic code for bootstrapping the central units 242b, 242m, and the volatile memories (RAM) 254b, 254m are used as working memories by the microprocessors 248b, 248m. The central units 242b, 242m are provided with mass memories 257b, 257m, for example solid state disks (SSD), for storing programs and data. Furthermore, the central unit 242b, 242m comprises a number of controllers 260b, 260m for peripherals or input / output (I / O) units. In particular, the controller 260b of the baby scope 115b controls the excitation light source 203b, the fluorescence camera 236b, the (possible) white light source 209b, the (possible) reflection camera 239b, the monitor 121b and the video interface 124b. Furthermore, the controller 260b can also control a setting device 263b used to set the corrections used to manually align the images of the baby scope 115b and the images of the mother scope 115m, for example a rotation / linear dial mounted on the probe or dolly of the baby scope 115b, a foot operated rotation / linear dial or a foot paddle, all of which can be operated either in a continuous mode or in a toggle mode, etc. In turn, the controller 260m of the MotherScope 115 controls the white light source 209m, the reflector camera 239m, the monitor 121m, and the video interface 124m.In both cases, controllers 260b, 260m can control further peripherals (not shown), such as a keyboard, a trackball, a drive for reading / writing a removable storage unit (such as a USB key), and a network interface card (NIC) for connecting to a (communications) network (such as a local area network (LAN)).

[0037] Generally, the field of view 221b of the baby scope 115b and the field of view 221m of the mother scope 115m are different. For example, the field of view 221b is smaller than the field of view 221m, and the fields of view 221b and 221m at least partially overlap (in the example shown in the figure, both the fields of view 221b and 221m include the target 109). In any case, the images of the baby scope 115b and the images of the mother scope 115m generally have different characteristics in terms of shape, zooming, resolution, contrast, noise, etc., and these different characteristics are much more evident between the fluorescent images of the baby scope 115b and the reflected images of the mother scope 115m due to these different properties. Furthermore, the probe of the baby scope 115b is generally not fixed in the working channel of the mother scope 115n (not shown) but is movable relative to it (along and in particular around its longitudinal axis). Inevitable (a priori unknown) movements of the acquisition unit of the baby scope 115b relative to the acquisition unit of the mother scope 115m usually occur (in combination with the flexible nature of the probes of the mother scope 115m and the baby scope 115b) due to, for example, the peristalsis, breathing and heartbeat of the patient, which results in corresponding misalignments (e.g. translations, especially rotations) between the images of the baby scope 115b and the mother scope 115m.

[0038] Referring now to FIG. 3, the general principle of the approach according to one embodiment of the present disclosure is illustrated.

[0039] In particular, Motherscope provided a series of reflectivity images (Motherscope images) of 305m taken over time, and the figure shows 305m0 and 305m -1 Only the last two, labeled with , are shown in the figure. Similarly, the BabyScope provides a corresponding series of fluorescence images and possibly reflectance images (BabyScope images) 305b. For simplicity, when only the fluorescence images are considered, the figure shows only the MotherScope images 305m0 and 305m1. -1 305b0 and 305b, which were obtained substantially simultaneously with -1 Shown are the last two Babyscope (fluorescence) images with the annotation:

[0040] In an approach according to an embodiment of the present disclosure, a computing device receives both the mother scope image 305m and the baby scope image 305b. For example, the central unit of the baby scope receives the mother scope image 305m transmitted from the mother scope and the baby scope image 305b acquired by the baby scope (or vice versa). Each (registration) pair of corresponding mother scope image 305m and baby scope image 305b is registered to make the mother scope image 305m and the baby scope image 305b spatially corresponding. For example, the mother scope image 305m and the baby scope image 305b are registered to correct misalignment caused by the movement between the baby scope probe and the mother scope probe. This operation can be performed manually or automatically (described in more detail below). A representation of the body part based on each pair of thus aligned Motherscope image 305m and Babyscope image 305b is then output, and, for example, the (aligned) Motherscope image 305m and Babyscope image 305b are displayed superimposed on a monitor connected to the computing device (in this case the Babyscope monitor).

[0041] As a result, it is possible to combine the fluorescent image of the Babyscope (which provides an accurate representation of the target of the medical procedure) with the reflectance image of the Motherscope (which provides a high-quality representation of the area of ​​interest of the medical procedure), which significantly facilitates the doctor's task and improves the outcome of the medical procedure with beneficial effects on the patient's health.

[0042] In a specific implementation of the technique according to an embodiment of the present disclosure, for each registered pair of a mother scope image 305m and a baby scope image 305b, the mother scope (mother scope) motion is estimated according to the mother scope image 305m, and the baby scope (baby scope) motion is estimated according to the baby scope image 305b, which are independent of each other, and the mother scope motion is estimated according to multiple mother scope images 305m corresponding to the registered ones (e.g., for the mother scope image 305m0, the mother scope image 305m0, the baby scope image 305b ... -1 ), the movement of the BabyScope is reflected by multiple BabyScope images 305 that correspond to the one being aligned. b (For example, for the baby scope image 305b0, the baby scope image 305b0, 305b -1 For example, for a simple translation as shown, the mother scope motion and baby scope motion (direction and magnitude) are defined by vector 310m and vector 310b, respectively.

[0043] Then, the mother scope image 305m0 and the baby scope image 305b0 are aligned according to the mother scope motion 310m and the baby scope motion 310b. For example, misalignment due to relative motion between the mother scope and the baby scope is determined by comparing the mother scope motion 310m and the baby scope motion 310b. In the example in question, the relative rotation of the baby scope with respect to the mother scope is defined by an angle 310mb equal to the difference between the vectors 310b and 310m, and a correction 310bm given by the opposite angle of this angle 310mb is applied to the baby scope image 305bo (by rotating it) to produce a corresponding (aligned) baby scope image 305b0 that is better aligned with the corresponding mother scope image 305m0. 0R is obtained.

[0044] In this way, the baby scope image and the mother scope image are continuously corrected, and finally, the effect of the relative movement between the baby scope and the mother scope is removed. As a result, the baby scope image and the mother scope image are aligned independently of the starting condition. In fact, the apparent movement of the fields of view of the mother scope and the baby scope helps to estimate their relative positions. For example, in the situation shown in the figure, the translation of the baby scope with respect to the mother scope gradually corrects their relative rotation.

[0045] More generally, the alignment of the motherscope image 305m and the babyscope image 305b may include any other transformation that brings them into spatial correspondence (i.e., references them to a common reference space), such as translation, rotation, zooming, etc. In particular (not shown), it is possible to determine the corresponding rotation center points of the motherscope and the babyscope, and then calculate the relative translation of the babyscope with respect to the motherscope, defined by the distance from the rotation center point of the motherscope to the rotation center point of the babyscope, and the babyscope image is translated by the inverse of the resulting vector.

[0046] As yet another example (also not shown), it is possible to determine the corresponding zoom vector fields of the mother scope and the baby scope, and then calculate a magnification difference (zoom) of the baby scope relative to the mother scope, defined by the difference between the zoom vector field of the mother scope and the zoom vector field of the baby scope, and the baby scope image is scaled by the inverse of the resulting magnification difference.

[0047] The above described technique allows the automatic alignment of baby scope images and mother scope images in an inherently robust manner. In fact, the proposed technique does not involve a comparison between baby scope images and mother scope images. Thus, a valid alignment is possible regardless of the fact that baby scope / mother scope images have very limited and not easily distinguishable features (as they mainly contain smooth patches with low contrast that are difficult to match in different images with known feature-based methods). Moreover, the various features of the baby scope images and mother scope images are completely irrelevant to the results obtained. A valid alignment is possible even when the baby scope images and mother scope images are significantly different, especially when the baby scope only provides a fluorescent image. In other words, the desired result is achieved indirectly by operating the baby scope images and mother scope images separately, rather than directly aligning the baby scope images and mother scope images. To this end, the inevitable relative movement of objects in the BabyScope's and MotherScope's fields of view is converted into an advantage; thus, the very cause of misalignment between the BabyScope and MotherScope images (the inherent movement between the BabyScope and MotherScope) is exploited to eliminate (or at least reduce) the misalignment.

[0048] Referring now to FIG. 4, there is shown the major software components that can be used to implement the technique according to one embodiment of the present disclosure.

[0049] All software components (programs and data) are generally indicated by the reference number 400. In a particular implementation, the software components 400 are stored in a mass memory and, when the programs run, are loaded (at least partially) into the working memory of the central unit of the BabyScope together with the operating system and other application programs that are not directly related to the disclosed approach (and therefore omitted in the figure for simplicity). The programs are initially installed in the mass memory, for example from a removable storage unit or from a network. In this respect, each program may be a module, a segment or a piece of code, which comprises one or more executable instructions for implementing a specified logical function.

[0050] In particular, the acquirer 403 drives components of the baby scope dedicated to acquiring (baby scope) fluorescence images and possibly (baby scope) reflectance images of the baby scope's field of view, suitably illuminated for this purpose, during a medical procedure. The acquirer 403 writes to a (baby scope) fluorescence image repository (storage location) 406 and a (baby scope) reflectance image repository 409, which contain corresponding sequences of baby scope fluorescence images and baby scope reflectance images, respectively, acquired consecutively during a medical procedure. The baby scope fluorescence image repository 406 and the baby scope reflectance image repository 409 contain a corresponding entry for each pair of baby scope fluorescence images and baby scope reflectance images acquired at the same acquisition time. This entry stores a bitmap of the corresponding baby scope (fluorescence / reflectance) image, which is defined by a matrix of cells (e.g. 512 rows and 512 columns), each containing the (color) values ​​of a pixel, i.e. an elementary pixel, representing a corresponding location of the baby scope's field of view. Each pixel value of the BabyScope fluorescence image defines the brightness of the pixel as a function of the intensity of the fluorescent light emitted by that location, while each pixel value of the BabyScope reflectance image defines the brightness of the pixel as a function of the intensity of the visible light reflected by that location (e.g., from 0 to 256 for each of its RGB components). A video interface drive 412 drives the BabyScope video interface, and in particular, as far as this disclosure is concerned, the video interface drive 412 receives (from its video interface) the (MotherScope) reflectance images acquired by the MotherScope during the medical procedure. The video interface drive 412 writes to a (MotherScope) reflectance image repository 415, which contains a series of MotherScope reflectance images acquired consecutively during the medical procedure by the MotherScope substantially synchronously with the BabyScope fluorescence / reflectance images. As mentioned above, the MotherScope reflectance image repository 415 contains an entry for each MotherScope reflectance image.An entry stores a bitmap of the Motherscope reflectance image, which is defined by a matrix of cells (typically with a higher resolution, e.g. with 2048 rows and 2048 columns), each containing a pixel (color) value (representing the corresponding location in the Motherscope's field of view) and defining the pixel's brightness as a function of the intensity of the visible light reflected at that location (e.g. from 0 to 256 for each of its RGB components).

[0051] A preparer 418 prepares the (baby scope / mother scope) reflectance images by pre-processing them for the alignment operation. The preparer 418 reads the baby scope reflectance image repository 409 and the mother scope reflectance image repository 415 and writes them to the (baby scope) prepared image repository 421 and the (mother scope) prepared image repository 424. The baby scope prepared image repository 421 and the mother scope prepared image repository 424 contain an entry for each baby scope reflectance image in the corresponding repository 409 and each mother scope reflectance image in the corresponding repository 415, respectively. The entry stores the corresponding baby scope / mother scope prepared (reflectance) image. The baby scope / mother scope prepared image is formed by a matrix of cells (generally with a smaller size compared to the corresponding baby scope / mother scope reflectance image), each storing a corresponding pixel value (e.g., a grayscale value ranging from 0 for white to 256 for black). If a baby scope reflectance image is not available, the same operations described above are performed on the baby scope fluorescence image (in addition to the mother scope reflectance image) to obtain a corresponding baby scope prepared (fluorescence) image. Thus, in this case, preparer 418 reads baby scope fluorescence image repository 406 (shown in dashed lines) instead of baby scope reflectance image repository 409.

[0052] The corrector (collector) determines the correction (i.e., transformation) to be applied to the images provided by the baby scope and the mother scope and automatically aligns the images according to these independently estimated movements. The corrector can implement optical flow techniques, deep learning techniques, or a combination of both. A corrector based on optical flow techniques comprises an estimator 427 that estimates the optical flow of each baby scope / mother scope prepared image. The optical flow of a baby scope / mother scope prepared image represents the apparent movement of the contents of the baby scope / mother scope's field of view due to its motion. The estimator 427 reads the baby scope prepared image repository 421 and the mother scope prepared image repository 424 and writes them to the (baby scope) motion vector repository 430 and the (mother scope) motion vector repository 433. The baby scope motion vector repository 430 and the mother scope motion vector repository 433 contain an entry for each baby scope / mother scope prepared image (different from the first one) in the corresponding repository 421, 424. This entry stores (e.g., transforms) the corresponding baby scope / mother scope motion vectors that define the optical flow of the baby scope / mother scope prepared images. The corrector further comprises a calculator 436 that calculates the correction for each (aligned) pair of corresponding baby scope prepared images and mother scope prepared images according to the corresponding baby scope / mother scope motion vectors. The calculator 436 reads the baby scope motion vector repository 430 and the mother scope motion vector repository 433. Alternatively, the corrector based on deep learning techniques comprises a neural network 439 that directly determines the correction for each aligned pair of baby scope prepared images and mother scope prepared images. The neural network 439 reads the baby scope prepared image repository 421 and the mother scope prepared image repository 424. Alternatively, a setter 442 is used to manually set the corrections (applied to the images provided by the baby scope and mother scope to align the images).For this purpose, the setter 442 exposes a user interface offering one or more (software) dials (e.g., rotation / linear sliders, input boxes, hotkeys, etc. designed to be controlled by a mouse, keyboard and / or trackball) and / or drives for a setting device allowing input of corrections (e.g., rotation angles, translation directions and distances, zooming factors, etc.). The calculator 436, the neural network 439 and the setter 442 write to the correction repository 445. In particular, in the case of the calculator 436 and the neural network 439 (automatic alignment), the correction repository 445 contains an entry for each aligned pair of babyscope / motherscope prepared images, which entry stores the corresponding correction, whereas in the case of the setter 442 (manual alignment), the correction repository 445 stores the (common) corrections of the (original) babyscope / motherscope images (with each correction defined by a rotation angle in both cases in the example in question).

[0053] An aligner 448 aligns the images provided by the baby scope and the mother scope according to the corresponding corrections. In particular, the aligner 448 always aligns each (aligned) pair of corresponding baby scope fluorescence images and mother scope reflectance images according to the corresponding corrections. For this purpose, the aligner 448 reads the baby scope fluorescence image repository 406 and the correction repository 445 and writes them to a (baby scope) aligned fluorescence image repository 451. The baby scope aligned fluorescence image repository 451 contains an entry for each baby scope fluorescence image in the corresponding repository 406. The entry stores the corresponding (baby scope) aligned fluorescence image. The baby scope aligned fluorescence image is formed by a matrix of cells (with the same size as the baby scope fluorescence image), each of which stores a corresponding (color) pixel value. Moreover, if the corrector is based on both optical flow and deep learning techniques (to determine the corrections incrementally), the aligner 448 further (pre-)aligns each (aligned) pair of corresponding babyscope-prepared and motherscope-prepared images according to these (coarse) corrections. For this purpose, the aligner 448 further reads and writes to the babyscope-prepared image repository 421.

[0054] The display 454 drives the monitor of the BabyScope to display a representation of the body part based on the BabyScope registered fluorescence image and the MotherScope reflectance image (e.g., superimposed on one another). The display 454 reads the BabyScope registered fluorescence image repository 451 and the MotherScope reflectance image repository 415.

[0055] 5, an operational diagram is shown illustrating the flow of activities associated with implementing a technique according to one embodiment of the present disclosure, in which each block may correspond to one or more executable instructions for implementing a given logical function on the BabyScope central unit.

[0056] In particular, the operational diagram depicts an exemplary process that may be used to image a patient during an endoscopic medical procedure (endoscopic procedure) using method 500 as described above.

[0057] Prior to an endoscopic procedure, a medical professional (e.g., a nurse) administers a fluorescent agent to the patient. The fluorescent agent (e.g., indocyanine green, methylene blue, etc.) is configured to reach and become substantially fixed within a specific (biological) target, such as a tumor to be examined / resected / treated. This result can be achieved by using either a non-targeted fluorescent agent (configured to accumulate within the target without specific interaction, e.g., passive accumulation), or a targeted fluorescent agent (configured to attach to the target by specific interaction, e.g., achieved by incorporating target-specific ligands into the formulation of the fluorescent agent, based on chemical binding properties and / or physical structures that can interact with various tissues, vascular properties, metabolic properties, etc.). For example, the fluorescent agent is administered intravenously to the patient as a bolus (e.g., using a syringe). As a result, the fluorescent agent circulates within the patient's vascular system until it reaches and binds to the target. The remaining (unbound) fluorescent agent is instead removed from the blood pool. After a waiting period (e.g., from a few minutes to 24-72 hours) during which the fluorescent agent accumulates in the (potential) tumor and is washed out from the rest of the patient, the endoscopic procedure can begin. Therefore, if necessary, the physician anesthetizes the patient (fully or partially). In either case, the physician inserts the Motherscope probe, switched on by the (medical) operator, into the patient's body cavity (through its opening) until its tip reaches the area of ​​interest in the body part where the tumor is likely to be present. In response, the Motherscope's white light source illuminates its field of view and the Motherscope's reflectance camera continuously acquires Motherscope reflectance images of its field of view, which are displayed in real time on the Motherscope's monitor (not shown).

[0058] As far as the technique according to one embodiment of the present disclosure is concerned, at a certain point in the endoscopic procedure, the physician inserts the baby scope's probe into the working channel of the mother scope (via its working port) until its tip reaches the same region of interest of the body part (the tip of the mother scope's probe). The physician may decide at any time to switch on the baby scope and start the imaging process (with the corresponding alignment operation). For example, this may be done before inserting the baby scope into the working channel of the mother scope or after the tip of the baby scope's probe has reached the tip of the mother scope's probe. In either case, in response, the (imaging) process starts by proceeding from the black start circle 503 to block 506. At this point, the acquirer switches on the excitation light source and the white light source for illuminating the field of view of the baby scope. The flow of operations then branches into various operations that are performed simultaneously. In particular, the acquirer acquires a new baby scope fluorescence image in block 509 and adds it to the corresponding repository. If necessary, the acquirer acquires a new baby scope reflectance image in block 512 and adds it to the corresponding repository. Furthermore, the video interface (continuously receiving the Motherscope reflectance images transmitted from the Motherscope) adds the last one of these to the corresponding repository in block 515. In this way, the Babyscope fluorescence image and the Babyscope reflectance image (if available) are acquired almost simultaneously, providing different representations (in terms of fluorescence and visible light, respectively) of the same field of view of the Babyscope that are spatially coherent (i.e. there is a predictable correlation between their pixels up to complete identity). Furthermore, the Motherscope reflectance images can also be considered as having been acquired substantially simultaneously, apart from a phase shift between the acquisition speeds of the Babyscope and the Motherscope (which is negligible in practice).

[0059] The flow of operations rejoins at block 509, block 512, block 515 to block 518. At this point, the process branches according to the mode (manual / automatic) of the baby scope alignment operation (e.g., manual setting, default defined, or only available). In the case of the automatic mode of the alignment operation, the process further branches at block 521 according to the configuration of the baby scope. If the baby scope is configured to acquire both fluorescence and reflectance images, the preparer retrieves the (last) baby scope reflectance image that was just added to the corresponding repository at block 524. Conversely, if the baby scope is configured to acquire only fluorescence images, the preparer retrieves the (last) baby scope fluorescence image that was just added to the corresponding repository at block 527. In either case, the process proceeds to block 530, where the preparer retrieves the (last) mother scope reflectance image that was just added to the corresponding repository. The preparer prepares the just retrieved baby scope (reflectance or fluorescence) image and mother scope (reflectance) image for the corresponding alignment operation at block 533. For example, the preparer can reduce each BabyScope / MotherScope image to its central region by discarding its parts surrounding a central region defined by 15-25%, e.g. 20%, of its (outermost) pixels from each border of the BabyScope / MotherScope image. In this way only the most useful parts of the BabyScope / MotherScope image are taken into account, with a corresponding increase in accuracy and reduction in computation time of the alignment operation. Additionally or alternatively, the preparer can reduce the size of each (possibly reduced) BabyScope / MotherScope image, for example by reducing its size by 40-60%, e.g. by 50% (e.g. using a low-pass filter followed by subsampling). This reduces the computation time of the alignment operation (e.g. by a factor of 3-4) and at the same time does not adversely affect its accuracy, but rather slightly improves it (as it becomes easier to detect the smallest movements).Additionally or alternatively, the preparer can convert each (possibly reduced and / or downscaled) BabyScope / MotherScope image to grayscale by replacing the RGB components of each pixel value with a single grayscale component representing the corresponding light intensity (e.g. from 0 for black to 255 for white, etc.), e.g. the grayscale component is calculated as a weighted average of the RGB components (to maintain the perceptual brightness). This reduces the computation time of the alignment operation and at the same time increases its accuracy and reduces its sensitivity to environmental conditions (e.g. illumination, contrast, equipment, etc.). In either case, the preparer adds the BabyScope / MotherScope prepared images thus acquired to the corresponding repository. After a transition time necessary to have a sufficient number of BabyScope / MotherScope prepared images for the alignment operation (e.g. 2-5), the (alignment) pair of BabyScope fluorescence image and MotherScope reflectance image (first acquired) is aligned. For this purpose, the flow of operations branches in block 536 according to the BabyScope configuration (e.g. set manually, defined by default, or the only one available).

[0060] If the alignment is based only or preliminary on deep learning techniques, a neural network is used to determine the corresponding corrections. Essentially, a neural network is a data processing system that approximates the operation of the human brain. A neural network includes basic processing elements (neurons), which perform operations based on corresponding weights. The neurons are connected through unidirectional channels (synapses), which transfer data between them. The neurons are organized into layers that perform various operations, always including an input layer that receives input data for the neural network and an output layer that provides output data. In one embodiment of the present disclosure, the neural network is a convolutional neural network (CNN), i.e., a specific type of deep neural network (one or more hidden layers are successively arranged between the input layer and the output layer along the processing direction of the neural network), one or more of which perform (cross)convolution operations. For example, the neural network is based on a modified version of the VGG-16 model. More specifically, the input layer is configured to receive an input image generated by concatenating an (estimated) set of BabyScope prepared images and an (estimated) set of prepared MotherScope images. Each estimation set includes multiple last BabyScope / MotherScope prepared images (e.g., the last 2-10), and the input image has the same size as the BabyScope / MotherScope prepared images and the number of channels (values ​​of each cell) given by the components of the pixel values ​​of the BabyScope / MotherScope prepared images (e.g., 4 channels for 2 BabyScope prepared images and 2 MotherScope prepared images, each with one grayscale component for each pixel value). The hidden layer first includes 5 groups each of one or more convolution layers (e.g., a first group of 2 convolution layers, a second group of 2 convolution layers, a third group of 3 convolution layers, a fourth group of 3 convolution layers, and a fifth group of 3 convolution layers), followed by a corresponding max pooling layer.In general, each convolutional layer performs a convolution operation through a convolution matrix (filter or kernel) defined by the corresponding weights (the same number between the corresponding apply data and the filter for each channel). This convolution operation is performed successively on a limited part of the apply data (receptive field) by shifting the filter across the apply data by a selected number of cells (stride), possibly adding cells with zero content around the border of the apply data (padding), which can also be filtered (adding a bias value to it, transforming it, and applying an activation function to introduce a nonlinear coefficient) to obtain the corresponding filtered data. Each element of the filtered data of the convolutional layer can also be interpreted as the output of a neuron, looking only at a small region (receptive field) in the apply data and sharing parameters with other neurons spatially to the left and right (according to the stride of the filter shared between all neurons of the convolutional layer). In the VGG-16 model, the convolutional layer applies a very small 3 × 3 filter with padding 1 and stride 1 (to maintain the size of the apply data), each of whose neurons applies a rectified linear unit (ReLU) activation function. Each max pooling layer is a pooling layer (which downsamples its application data to reduce sensitivity to differences due to small displacements) that replaces the values ​​of each limited portion (window) of the application data with their maximum value by shifting the window over the application data by a selected number of cells (stride). In the VGG-16 model, the max pooling layer has a 2×2 window with a stride of 2. Overall, the group of hidden layers described above reduces the size of the application data to 1×1 and increases these channels to 4096. The hidden layers then comprise three fully connected (or dense) layers. Each fully connected layer has neurons connected to all neurons in the previous layer. The fully connected layers progressively reduce the number of channels to a number corresponding to all possible classes of the input data. In the VGG-16 model, the fully connected layers apply a ReLU activation function and reduce the channels to 1000 in the ILSVRC classification.In one embodiment of the present disclosure, the channels are reduced to provide 360 ​​channels for the degree of misalignment rotation angle (from 0° to 359°) between the (registered) pair of the last BabyScope prepared image and the last MotherScope prepared image. This is a good compromise between the conflicting requirements of low complexity of the neural network (affecting training and response time) and high resolution (affecting accuracy). Furthermore, in the first two fully connected layers, the ReLU activation function is replaced by a linear activation function. Finally, the hidden layer comprises a softmax layer, which normalizes the application data to a probability distribution (each value ranges from 0 to 1, and the sum of all values ​​is 1). Then, the output layer determines the correction for the registered pair, which is set to the opposite of the rotation angle with the highest probability. In this deep learning technology-based implementation, the process proceeds from block 536 to block 539, where the estimated set of BabyScope images and the estimated set of MotherScope images are concatenated to the input image of the neural network and applied to it. Moving to block 542, the neural network directly outputs a correction for the corresponding registration pair of the babyscope / motherscope prepared images. This correction is added to the value present in the last entry of the corresponding repository (initialized to 0 and possibly preliminary set using the optical flow technique described below). As a result, if only the implementation based on deep learning techniques is applied, the correction determined therein directly defines the final value of the registration operation. Conversely, if the implementation based on optical flow techniques was previously applied, the correction determined by the deep learning technique improves the preliminary value determined by the optical flow technique. The flow of operation further branches in block 545 according to the configuration of the babyscope. If the correction thus obtained is the final value to be applied to the registration pair of the babyscope fluorescence image and the motherscope reflectance image as initially acquired (the implementation based on optical flow techniques is not applied later), the aligner aligns them accordingly in block 548.For example, the aligner applies a correction to the baby scope fluorescence image (by rotating it in the example in question) and adds the baby scope aligned fluorescence image thus obtained to the corresponding repository. Since the implementation based on deep learning technology is extremely fast, it is well suited for real-time applications in which the baby scope fluorescence image and the mother scope reflection image are aligned with a short delay from their acquisition so that their display during the endoscopic procedure is possible. Furthermore, if the implementation based on deep learning technology is applied after the implementation based on optical flow technology, the neural network can be configured to provide a much smaller number of rotation angle values ​​(e.g., 10 values ​​from 0° to 9°). In this case, the neural network can be trained more extensively and will be much more accurate. Referring back to block 545, if the correction thus obtained is still a preliminary value to be adjusted by applying the optical flow technology, the aligner aligns the aligned pair of prepared baby scope / mother scope images in block 551 by similarly applying the correction to the prepared baby scope image and replacing it in the corresponding repository. The process then proceeds to block 554.

[0061] If the alignment is based solely or preliminarily on optical flow techniques, the same point is reached directly from block 536. In both cases, the estimator determines the baby scope motion vector (e.g. translation) of an (estimated) set of baby scope prepared images comprising a number of last baby scope prepared images, for example the last baby scope prepared image with respect to the penultimate baby scope prepared image (extracted from the corresponding repository), and adds the value thus obtained to the corresponding repository. Similarly, the estimator determines in block 557 the mother scope motion vector (e.g. translation again) of an (estimated) set of mother scope prepared images comprising a number of last mother scope prepared images, for example the last mother scope prepared image with respect to the penultimate mother scope prepared image (extracted from the corresponding repository), and adds the value thus obtained to the corresponding repository. For this purpose, the estimator can apply any known technique, for example being able to calculate local motion vectors at the level of one or more pixel groups (for example using a block matching algorithm, or an estimation of dense or sparse vector fields, etc.). Each of these represents the offset of the corresponding pixel group from the penultimate babyscope / motherscope prepared image to the last babyscope / motherscope prepared image. A (global) motion vector is then determined according to these local motion vectors (e.g., set to their average). The calculator calculates in block 560 the misalignment (in the example in question, the rotation angle) between the aligned pair of the last prepared babyscope / motherscope images according to the corresponding motion vector (e.g., equal to the angle from the motion vector of the babyscope prepared image to the motion vector of the motherscope prepared image) and adds it to the corresponding repository. The calculator calculates in block 563 a correction for the aligned pair of babyscope / motherscope prepared images according to a (calculated) set of a number of last misalignments (e.g., the last 2 to 10). In particular, the correction is calculated as the inverse of the weighted average of the misalignments of the calculated set.For example, a weight is assigned to each misalignment that decreases (for example exponentially) with its age and / or degree. A weight that decreases with age smooths the correction (by reducing the effects of jitter), whereas a weight that decreases with degree improves accuracy (by reducing the effects of misleading no or very small movements). The calculator then adds the correction thus obtained to the value present in the last entry of the corresponding repository (initialized to 0 and, if possible, to one previously set using deep learning techniques as described above). As a result, if only an implementation based on optical flow techniques is applied, the correction determined using it directly defines the final value for the alignment operation. Conversely, if an implementation based on deep learning techniques was previously applied, the correction determined using optical flow techniques adjusts its preliminary value determined using deep learning techniques. The flow of operations further branches in block 566 according to the configuration of the baby scope. If the correction thus obtained is the final value to be applied to the aligned pair of BabyScope fluorescence image and MotherScope reflectance image as originally acquired (the implementation based on deep learning techniques is not applied later), the aligner aligns the aligned pair of BabyScope fluorescence image and MotherScope reflectance image (as originally acquired) by applying the correction to the BabyScope fluorescence image in block 569 and adding the BabyScope aligned fluorescence image thus obtained to the corresponding repository. The implementation based on motion flow techniques is extremely accurate and therefore it is well suited for offline applications (where the required computation time is not critical) or for real-time applications for adjusting the preliminary alignment provided by the implementation based on deep learning techniques (where the latter reduces the computation time significantly).Returning to block 566, if the corrections thus obtained are still preliminary values ​​to be adjusted by applying deep learning techniques, the aligner similarly aligns the prepared babyscope / motherscope image registration pair by applying the corrections to the prepared babyscope image and replacing it in the corresponding repository in block 572. The process then returns to block 539.

[0062] Returning to block 518, in the case of a manual mode of the alignment operation, the setter verifies in block 575 whether the corrections (applied to the first acquired alignment pair of babyscope / motherscope images) need to be set. In particular, this always occurs at the beginning of the endoscopic procedure (to initialize the corrections) and possibly during the time when the physician determines that the corrections are no longer accurate (to update the corrections). If the corrections need to be set, the process proceeds to block 578, where the physician inputs the corrections via the setter. For example, the display continuously displays both the motherscope reflectance image and the babyscope reflectance image on the monitor of the babyscope (if the babyscope reflectance image is not available, the same operation is performed with the babyscope fluorescence image). The physician waits until a (relatively) stable state of the endoscopic procedure is reached (as indicated by the displayed motherscope / babyscope reflectance image). For example, this occurs when the tip of the motherscope probe reaches the area of ​​interest of the body part and the babyscope probe is inserted into the working channel of the motherscope with its tip reaching the same area of ​​interest of the body part. The physician now acts on the Mother / Baby Scope reflectance images (via the user interface or configuration device) until they appear to be aligned. This action can be performed on the Mother / Baby Scope reflectance images while they are being played back during their acquisition or paused depending on a corresponding command entered by the physician into the Baby Scope (for example, using its keyboard). Once the action is completed, the physician confirms the (manual) alignment of the Mother / Baby Scope reflectance images thus obtained by entering a corresponding command into the Baby Scope (for example, using its keyboard). In response, the configurer determines in block 581 the corrections corresponding to this manual alignment and stores them in the corresponding repository (replacing the previous values ​​initialized to null). The process then proceeds to block 584. The same point is reached directly from block 575 if no setting of corrections is required.At this point, the aligner aligns the aligned pair of BabyScope fluorescence image and MotherScope reflectance image (the one acquired first) by applying this correction (extracted from the corresponding repository) to the BabyScope fluorescence image and adding the resulting BabyScope aligned fluorescence image to the corresponding repository.

[0063] The flow of operations then rejoins in block 587 from block 548 (automatic mode based solely or additionally on deep learning techniques), from block 569 (automatic mode based solely or additionally on optical flow techniques), or from block 584 (manual mode). At this point, the display may display a representation of the body part based on the MotherScope reflectance image and the BabyScope registered fluorescence image. For example, the display may display an overlay image generated by overlaying the BabyScope registered fluorescence image on the MotherScope reflectance image (each pixel value of the overlay image is equal to the corresponding pixel value of the BabyScope registered fluorescence image if its brightness is higher (preferably strictly) than a threshold value (e.g., 5-10% of the maximum value) or is otherwise equal to the corresponding pixel value of the MotherScope reflectance image).

[0064] Referring now to block 590, if the imaging process is still in progress, the flow of operations loops back to before blocks 509-515 to continuously repeat the same operations. Conversely, if the imaging process is terminated, as indicated by an end command entered by the operator into the BabyScope (e.g., using its keyboard), the process ends at the concentric white / black stop circles 593 (after turning off the excitation light source and the white light source by the acquirer).

[0065] Referring now to FIG. 6, there is shown a schematic block diagram of a computer system 600 that can be used to train a neural network in a technique according to one embodiment of the present disclosure (training).

[0066] The training (computer) system 600 (e.g., a personal computer (PC)) comprises several units interconnected via a bus structure 605. In particular, a microprocessor (μP) 610 or more provides the logic capabilities of the training system 600. A non-volatile memory (ROM) 615 stores the basic code for bootstrapping the training system 600, and a volatile memory (RAM) 620 is used by the microprocessor 610 as a working memory. The training system 600 is provided with a mass memory 625, for example a solid-state disk (SSD), for storing programs and data. Furthermore, the training system 600 comprises several controllers 630 for peripherals or input / output (I / O) units. For example, the peripherals include a keyboard, a mouse, a monitor, a network adapter (NIC) for connecting to a communication network (e.g., the Internet), a drive for reading and writing removable storage units (e.g., USB type).

[0067] Referring now to FIG. 7, there is shown the main software components that can be used to train a neural network in a technique according to one embodiment of the present disclosure.

[0068] All software components (programs and data) are generally designated by the reference numeral 700. The software components 700 are stored in mass memory and, when the programs are executed, are loaded (at least partially) into the working memory of the training system together with the operating system and other application programs not directly related to the disclosed techniques (omitted from the figure for simplicity). The programs are initially installed into the mass memory, for example, from a removable storage unit or a network. In this regard, each program may be a module, segment or portion of code that includes one or more executable instructions for implementing certain logical functions.

[0069] The loader 705 loads a number of sample motherscope (reflectance) images acquired as described above during a sample endoscopic procedure or the like, together with a corresponding sample motherscope (e.g., by imaging the colon of a patient with a tumor). If possible, the (sample) motherscope images are acquired using various models of sample motherscopes to desensitize the neural network to the actual model of the motherscope used in reality. The loader 705 writes to a (sample) motherscope image repository 710 storing the motherscope images. A compositor 715 synthesizes a corresponding (babyscope) composite image from the motherscope images. This composite image mimics a corresponding reflectance image acquired by a babyscope or the like inserted via the working channel of the sample motherscope. The compositor 715 reads the motherscope image repository 710 and writes to a composite image repository 720 storing the composite image. A generator 725 generates a number of (sample) babyscope (reflectance) images from the composite image by applying various misalignments (e.g., multiple randomly selected rotations to each image). The generator 725 reads the composite image repository 720 and writes to the (sample) babyscope image repository 730. The babyscope image repository 730 contains an entry for each babyscope image. This entry stores the babyscope image, an indication of the motherscope image used to composite the corresponding composite image (e.g., its index in the corresponding repository 710), and a (reference) correction equal to the inverse of the misalignment used to generate the babyscope image (representing its gold value). A trainer 735 trains the neural network 439. The trainer 735 reads the motherscope image repository 710 and the babyscope image repository 730. In addition, the trainer 735 runs (applies input data and receives output data) a copy of the neural network, also indicated by reference number 439, and writes its weights.

[0070] Referring now to FIG. 8, an operational diagram illustrating the flow of operations associated with training a neural network in a technique according to one embodiment of the present disclosure is shown.

[0071] In particular, the operational diagram depicts an example process that can be used to train a neural network (during BabyScope's development and possibly during its next maintenance) using method 800. As noted above, each block corresponds to one or more executable instructions for implementing a particular logical function on the training system.

[0072] The process starts at a black start circle 805 and proceeds to block 810, where the trainer loads the Motherscope images into a corresponding repository (e.g., via a removable storage unit or a network). The compositor synthesizes a corresponding composite image from each Motherscope image (retrieved from the corresponding repository) in block 815. To this end, the compositor applies multiple updates to the Motherscope image with the aim of making it resemble a real reflectance image acquired using a Babyscope. In particular, the Motherscope image (which generally has a polygonal shape, such as a rectangle or an octagon) is converted to a circular shape (e.g., by resetting to zero the values ​​of pixels that are outside a small circle inscribed in the Motherscope image). Additionally or alternatively, the resolution of the Motherscope image is reduced (e.g., by downsampling it by 30-70%, e.g., by 50%). Additionally or alternatively, the contrast of the Motherscope is reduced (e.g., by applying a low-contrast filter). Additionally or alternatively, the Motherscope image is zoomed in to simulate a smaller field of view (for example, by duplicating the values ​​of the pixels in the central zone). Additionally or alternatively, the Motherscope image is transformed (for example, by random values). Additionally or alternatively, noise is added to the Motherscope image (for example, by adding random noise to its pixel values). The combiner adds the thus obtained composite image to the corresponding repository (associating it with the same position of the corresponding Motherscope image in that repository). The generator generates, in block 820, babyscope images from the composite images. In particular, for each composite image (retrieved from the corresponding repository), the generator generates a number of pseudorandom values ​​of the rotation angle (for example, between 100 and 400) and applies these to the composite image. The generator adds each thus obtained babyscope image to the corresponding repository together with the indication of the corresponding Motherscope image (for example, equal to the common index of the composite / motherscope images in these repositories) and the corresponding reference correction (equal to the opposite value of the applied rotation angle).

[0073] The trainer performs a training operation of the neural network to find (optimized) values ​​of its weights and, if possible, its parameters to optimize performance. For this purpose, the trainer selects, in block 825, several training sets, each formed by a babyscope image, a corresponding motherscope image, and a corresponding reference correction. The training sets are defined by sampling the babyscope images and selecting a percentage of them (e.g., 50% selected randomly). The trainer randomly initializes, in block 830, the weights of the neural network. Then, in block 835, the trainer enters a loop in which the trainer feeds the babyscope images of each training set successively (in any order) to the neural network and obtains the corresponding (estimated) correction. The trainer calculates, in block 840, a loss value based on the difference between the estimated correction of the training set and the reference correction (e.g., by applying a Huber function to limit the sensitivity to outliers). The trainer verifies, in block 845, whether the loss value is not acceptable and still improves significantly. This operation can be performed either in an iterative mode (after processing each training set for its loss value) or in a batch mode (after processing all training sets for the cumulative value of the loss value, e.g., its average). If so, the trainer updates the weights of the neural network in block 850 to improve the performance of the neural network. For example, in a process based on the Stochastic Gradient Descent (SGD) algorithm, the direction and amount of change is given by the gradient of a loss function that represents the loss value as a function of the weights, which is approximated using a backpropagation algorithm. The process then returns to block 835 and repeats the same operations. Referring again to block 845, the loop ends if the loss value becomes acceptable or if the weight changes do not provide a significant improvement (meaning that a minimum, at least a local or flat area, of the loss function has been found).The above loop can be performed by adding random noise to the weights, and / or it can be iterated starting from different initializations of the neural network, finding different (possibly better) local minima and identifying flat regions of the loss function.

[0074] Once the configuration of the neural network that provides the optimal minimum of the loss function is found, the process continues to block 855. At this point, the trainer performs a validation operation of the performance of the neural network thus obtained. For this purpose, the trainer selects multiple validation sets, each formed by a babyscope image, a corresponding motherscope image, and a corresponding reference correction. For example, a validation set is defined by babyscope images different from those of the training set. Then, in block 860, a loop is entered in which the trainer feeds the babyscope images of the (current) validation set (starting from the first one in any order) to the neural network and obtains the corresponding (estimated) correction. In block 865, the trainer calculates a loss value as described above based on the difference between the estimated correction of the validation set and the reference correction. In block 870, the trainer verifies whether the last validation set has been processed. If not, the flow of operations returns to block 860 and repeats the same operations for the next validation set. Conversely (when all validation sets have been processed), the loop ends by proceeding to block 875. At this point, the trainer determines the global loss of the validation (e.g., equal to the average of the loss values ​​of all validation sets). The flow of operations branches in block 880 according to the global loss. If the global loss is higher than the tolerance value (as strictly as possible), this means that the generalization ability of the neural network (from the configuration learned from the training set to the validation set) is too low. In this case, the process returns to block 825 and repeats the same operations with a different training set, training parameters (e.g., learning rate, number of epochs, etc.) and / or neural network parameters (e.g., number of offsets, activation function, etc.), or proceeds to block 820 to increase the number of baby scope images (not shown). Conversely, if the global loss is lower than the tolerance value (as strictly as possible), this means that the generalization ability of the neural network is satisfactory. In this case, the trainer accepts the configuration of the neural network for deployment to a batch of baby scope instances in block 885. The process then ends with concentric white / black stop circles 890 .

[0075] 9A-9B, various application examples of the technique according to one embodiment of the present disclosure are shown.

[0076] Starting from FIG. 9A, this concerns an exemplary alignment during an endoscopic procedure. In particular, the figure shows an aligned pair of baby scope (reflectance) image 905b and mother scope (reflectance) image 905m with the addition of a corresponding local motion vector. For baby scope image 905b and mother scope image 905m, a baby scope motion vector 910b and a mother scope motion vector 910m are determined, respectively (from the local motion vectors). A correction (rotation angle) 910bm is calculated as the mother scope motion vector 910m minus the baby scope motion vector 910b. Then, the baby scope image 905b is aligned (rotated by the corresponding angle) according to the correction 910bm, and as can be seen, a corresponding baby scope aligned (fluorescence) image 910br is obtained. Here, the baby scope aligned image 910br is substantially aligned with the mother scope image 905m (shown again without the local motion vector).

[0077] Turning now to Figure 9B, which concerns an exemplary training of a neural network. In particular, this figure shows two Motherscope (reflectance) images 915m, 920m acquired during a sample endoscopic procedure. Two synthetic (reflectance) images 915s, 920s (mimicking the corresponding Babyscope reflectance images) are synthesized from the Motherscope images 915m, 920m, respectively (by converting them to a circular shape, reducing the resolution / contrast, and zooming in).

[0078] (Modification) In order to meet local and specific requirements, those skilled in the art may apply many logical and / or physical modifications and changes to the present disclosure. More specifically, although the present disclosure has been described with reference to one or more embodiments thereof in a certain degree of detail, it should be understood that various omissions, substitutions and changes in form, details and other embodiments are possible. In particular, various embodiments of the present disclosure may be practiced without the specific details (e.g., numerical values) set forth in the foregoing description to provide a more complete understanding, but conversely, well-known features may be omitted or simplified so as not to obscure the description in unnecessary detail. Furthermore, it is expressly intended that a specific element and / or method step described in connection with any embodiment of the present disclosure may be incorporated in any other embodiment as a matter of general design choice. Moreover, items presented in the same group and various embodiments, examples or alternatives should not be construed as being equivalent to each other in nature (but are separate and autonomous entities). In all cases, each numerical value should be read as modified according to the applicable tolerance, and unless otherwise indicated, the terms "substantially," "about," and "approximately" should be understood to mean within 10%, preferably within 5%, and even more preferably within 1%. Moreover, each numerical range should be intended as explicitly specifying any possible number along a continuum within the range (including its endpoints). Ordinal numbers or other modifiers are merely used as labels to distinguish between elements with the same name and do not, in themselves, imply a priority, preference, or order.The terms "include," "comprise," "have," "contain," "involve," etc. are intended to have an open, non-exhaustive meaning (i.e., not limited to the listed items); the terms "based on," "dependent on," "according to," "function of," etc. are intended to have a non-exhaustive relationship (i.e., including possible additional variables); the term "a / an" is intended to refer to one or more items (unless expressly stated otherwise); and the term "means for" (or any means-plus-function form) is intended to refer to any structure adapted or configured to perform the relevant function.

[0079] For example, one embodiment provides a method for imaging a body part of a patient using an endoscope system. However, this (imaging) method can be used to image any body part (e.g., one or more organs, regions, tissues, such as the gastrointestinal tract, airway, urinary tract, uterus, internal joints, etc.) and of any patient (e.g., human, animal, etc.) using an endoscope system (see below). Furthermore, the method can be used for any medical procedure (e.g., surgery, diagnosis, treatment, etc.). The method only concerns operations for controlling the endoscope system, independent of interaction with the patient (or at most without substantial physical intervention on the patient that would require skilled medical expertise or involve health risks for the patient). In any case, the method can facilitate the task of the physician, but only by providing intermediate results that can assist them, strictly with the medical act always performed by the physician himself.

[0080] In one embodiment, the method includes acquiring (with a first endoscope unit of an endoscopic system) a first sequence of a plurality of first images of a first field of view including at least a portion of the body part. However, the first images may be of any number and type (e.g., reflectance images, luminescence images, ultrasound images, etc. representing the entire body part or only a portion thereof), which may be acquired with any first endoscopic unit (see below).

[0081] In one embodiment, the first sequence of first images is acquired via a first probe of a first endoscope unit, however the first probe can be of any type (e.g. based on a fiber bundle, chip-on tip, etc.).

[0082] In one embodiment, the method includes acquiring (with a second endoscope unit of an endoscope system) a second sequence of a plurality of second image sets corresponding to the first images, although the first and second image sets can be acquired in any manner (e.g., by acquiring them independently with different start times and frequencies and correlating each image of a sequence with the last image of the other sequence available or with previous or next images of the other sequence, by synchronously acquiring the first and second images, by receiving the images of one sequence, determining their phase and frequency, and synchronously starting acquisition of the images of the other sequence, etc.), and the second images can be acquired with any second endoscope unit (see below).

[0083] In one embodiment, each of the second image sets includes one or more second images of a second field of view that includes at least a portion of the body part. However, each second image set may include any number of second images of any type (e.g., luminescence images, reflectance images, ultrasound images, any combination thereof, representing the entire body part or only a portion thereof). The first and second fields of view may be of any type (e.g., different from each other, e.g., overlapping to some extent or discrete, equal to each other, e.g., for imaging a thinner hollow organ that branches off from a thicker hollow organ, as in cholangioscopy).

[0084] In one embodiment, the second sequence of second images is acquired via a second probe of a second endoscope unit that is movable relative to the first probe. However, the second probe may be of any type (e.g., based on chip-on-chip, fiber bundles, etc.). Furthermore, the two probes may be movable between them in any manner that is not known a priori (e.g., rotationally, translating, any combination of these, always or only before coupling, etc.).

[0085] In one embodiment, the method includes registering (by a computing device) each registered pair of a first image of the plurality of first images and a corresponding second image of the second image set. However, the registration may be of any type (e.g., affine transformations such as rotation, translation, both, non-rigid transformations such as zooming, warping, etc.). Furthermore, the registration may be performed by any computing device (see below), in any manner (e.g., automatically, semi-automatically, or manually, by updating only the first image, only the second image, or both), at any time (e.g., by determining a correction that is applied continuously if the two probes are always movable between them, only when the two probes are joined if they are subsequently fixed relative to each other, etc.).

[0086] In one embodiment, the method includes outputting (to an output unit) a representation of the body part based on the first and second images of each registered pair that are registered. However, the representation of the body part may be output to any output unit (see below) and in any manner (e.g., displayed, printed, transmitted remotely, in real time or offline, etc.). Furthermore, the representation of the body part may be based on the registered images in any manner (e.g., by outputting the images of each registered pair overlaid or side-by-side).

[0087] Further embodiments provide additional advantageous features, however these may be omitted altogether in the basic implementation.

[0088] In particular, in one embodiment, the method includes acquiring a second sequence of a second image set via a second probe removably inserted into the working channel of the first endoscope unit (using the second endoscope unit). However, the second probe can be inserted into the working channel in any manner (e.g., after or before inserting the motherscope into the cavity, etc.). In any case, the possibility of positioning the motherscope and the babyscope in any other manner (e.g., inserted independently into the cavity, the babyscope inserted into a trocar port of the laparoscope, etc.) is not excluded.

[0089] In one embodiment, the computing device is included between the first endoscope unit and the second endoscope unit, however, the computing device may be included in either one of the endoscope units.

[0090] In one embodiment, the method includes receiving (by a computing device) a corresponding first sequence of first images or a second sequence of a second set of images from the other of the first and second endoscopic units, although these (first or second) images may be received in any manner (e.g., in push mode, pull mode, via any wired / wireless communication channel, from a removable storage unit, downloaded from a network, etc.).

[0091] In one embodiment, the method includes determining (by a computing device) at least one correction for each registered pair of the first and second images due to corresponding relative motion of the first and second probes. However, the correction may be of any type (e.g., one correction value for the first or second image of the registered pair, two corresponding correction values ​​for the first and second images of the registered pair, etc.) and it may be determined in any manner (e.g., using only optical flow techniques, using only deep learning techniques, using both of these in any order, manually, etc.).

[0092] In one embodiment, the method includes aligning (by a computing device) each registered pair of a first image and a second image according to a corresponding correction, however, the registered pair of a first image and a second image may be aligned according to the correction in any manner (e.g., applying it completely, applying it in stages, etc.).

[0093] In one embodiment, the method includes estimating (by a computing device) a first motion of the first probe and a second motion of the second probe for each current image of the plurality of first images and for each current image of the plurality of second image sets, respectively, independently, where the first motion is estimated according to an estimated set of the plurality of first images corresponding to the current first image, and the second motion is estimated according to an estimated set of the plurality of second images corresponding to the current second image. However, the motion may be of any type (e.g., rotation, translation, both, etc.) and in any manner (e.g., motion vectors, rotation center points, zooming vector fields, etc.). The motion may be estimated according to any type of estimated set of first / second images (e.g., including any number of images, selected among all previous images, or with some temporal subsampling, etc.), and in any manner (e.g., optical flow techniques, deep learning techniques, any combination thereof, etc.).

[0094] In one embodiment, the method includes determining (by a computing device) a correction for each aligned pair of first and second images according to a first and second motion of one or more calculation sets of the plurality of first images and the plurality of second images, respectively, corresponding to the pair of first and second images. However, the correction may be determined in any manner (e.g., by calculating a misalignment between each pair of corresponding first and second images and calculating the correction from the misalignment in the calculation set, calculating a first and second displacement from the first and second motions in the calculation set, respectively, and calculating the correction from the first and second displacements, etc.) according to the first / second motions of the calculation set (e.g., as an average, median, mode, etc. of associated values, weighted in any manner or as is).

[0095] In one embodiment, the method comprises estimating (by a computing device) a first motion vector indicative of a first motion for each current first image according to the estimated set of corresponding first images and a second motion vector indicative of a second motion for each current second image according to the estimated set of corresponding second images. However, the motion vectors may be of any type (e.g. a dense vector field with two components for each location, a sparse vector field with vector components at some of the locations, etc.) and they may be estimated in any way (e.g. directly at the level of the whole image, by applying methods such as block matching, phase correlation, differentiation, etc., by aggregating in any way the local values ​​for that location or groups thereof according to the average, dominant motion, etc.).

[0096] In one embodiment, the method includes calculating (by a computing device) a misalignment of each pair of corresponding current first and second images according to corresponding first and second movements, however the misalignment may be calculated in any manner (e.g., motion vector difference, rotation center distance, difference between zooming vector fields, etc.).

[0097] In one embodiment, the method includes calculating (by a computing device) a correction for each registered pair of first and second images according to the misalignment of the corresponding calculated set of first and second images, however, the correction may be calculated in any manner depending on the misalignment (e.g., as the mean, median, mode, etc. of the misalignment, weighted in any manner or raw).

[0098] In one embodiment, for this purpose, misalignments are weighted decreasingly with their corresponding range, however, misalignments may be weighted in any manner according to their range (e.g., according to a continuous or discrete function, exponentially, linearly, etc., by simply ignoring those below a threshold, etc.).

[0099] In one embodiment, for this purpose, misalignments are weighted decreasingly with the corresponding temporal distance from the aligned pair of the first and second images, however, misalignments may be weighted in any manner (e.g., the same or different manner to those described above) according to these temporal distances.

[0100] In one embodiment, the method includes providing (by a computing device) input data to a neural network based on a further set of estimates of the first image and a further set of estimates of the second image for each registered pair of the first and second images. However, the neural network may be of any type (e.g., a convolutional neural network, a time-delay neural network, a recurrent neural network, a modular neural network, etc.) and it may provide input data based on any number of first / second images in any manner (e.g., a common concatenation of the first and second images, a concatenation of the first image, a concatenation of the second image, a first / second image separately, etc.).

[0101] In one embodiment, the method includes receiving (by a computing device) a correction for each registered pair of the first and second images from a neural network, although the neural network may provide the correction in any manner (e.g., defining a classification result, a regression result, etc.).

[0102] In one embodiment, the method includes selecting (by a neural network) a correction for each registered pair of the first and second images from among a plurality of predefined corrections, although the predefined corrections may be any number and any type (e.g., uniformly distributed, more concentrated in a particular range, etc.).

[0103] In one embodiment, the neural network is a convolutional neural network, however, the convolutional neural network may be of any type (e.g., derived from a known model, e.g., VGG-16, VGG-19, ResNet50, etc., custom, etc.).

[0104] In one embodiment, the convolutional neural network includes multiple groups (along the processing direction of the convolutional neural network), each group including one or more convolutional layers followed by a max pooling layer. However, there may be any number of groups, each including any number and type of convolutional layers followed by any type of max pooling layer (e.g., with any receptive field, stride, padding, activation function, etc.).

[0105] In one embodiment, the convolutional neural network includes multiple fully connected layers that provide corresponding probabilities of predefined corrections (along the processing direction of the convolutional neural network). However, the fully connected layers may be of any number and type (e.g., with any number of channels, activation functions, etc.).

[0106] In one embodiment, the method includes acquiring (with a first endoscopic unit) a first sequence of first images, each of which includes a plurality of first values ​​representing corresponding locations of the first field of view, although each first value may include any number and type of components (e.g., RBG, XYZ, CMYK, grayscale, etc.).

[0107] In one embodiment, the method includes acquiring (with the second endoscopic unit) a second sequence of a second image set, each of the second image sets including a plurality of second values ​​that represent corresponding locations of the second field of view. However, each second value may include any number and type of components (e.g., the same or different from the first components).

[0108] In one embodiment, the method includes providing input data to the neural network for each registration pair of a first and second image, the input data being obtained by concatenating (by a computing device) for each location a first value of the first image of the further estimation set and a second value of the second image of the further estimation set, however, the images may be concatenated in any manner (e.g., the first image followed by the second image, vice versa, etc.).

[0109] In one embodiment, the method includes receiving a correction that is manually entered (by a computing device), however, the correction may be entered in any manner (e.g., via software and / or hardware commands that are partial, different, and / or additional to those described above).

[0110] In one embodiment, the method includes acquiring (with a first endoscope unit) a first sequence of first images that are corresponding reflectance images representative of visible light reflected by content in a first field of view, however the reflectance images may be of any type (e.g., color, black and white, hyperspectral, etc.).

[0111] In one embodiment, the method includes acquiring (with the second endoscope unit) a second sequence of a second set of images, which are corresponding luminescence images (representing luminescence light emitted by luminescent materials in the second field of view) and corresponding further reflectance images (representing visible light reflected by the contents of the second field of view). However, the luminescence light may be of any type (e.g., NIR, infrared (IR), visible light, etc.) and it may be emitted by any exogenous / intrinsic or exogenous / intrinsic luminescent material (e.g., any luminescent agent, any natural luminescent component, etc., based on any luminescence phenomenon, e.g., fluorescence, phosphorescence, chemiluminescence, bioluminescence, Raman radiation, etc.) in any manner (e.g., in response to a corresponding excitation light, or more generally any other excitation other than heating). Furthermore, the further reflectance image may be of any type (e.g., the same or different to the reflectance image).

[0112] In one embodiment, the method comprises a step of determining (by a computing device) a correction for each registered pair of reflectance and luminescence images according to the corresponding reflectance image and the further reflectance image, although the possibility of determining the correction according to the reflectance and luminescence images is not excluded.

[0113] In one embodiment, the method comprises a step of registering (by a computing device) each registered pair of a reflectance image and a luminescence image according to a corresponding correction, however, the possibility of registering each registered pair of a reflectance image and a further reflectance image is not excluded.

[0114] In one embodiment, the method comprises a step of outputting a representation of the body part based on the reflectance image and the luminance image of each registered pair (to an output unit), however, the possibility of outputting a representation of the body part based on further (registered) reflectance images is not excluded.

[0115] In one embodiment, the method includes acquiring (with the second endoscope unit) a second sequence of a second set of images that are corresponding luminescence images representative of luminescence light emitted by the luminescent material in the second field of view, however the luminescence images may be of any type (see above).

[0116] In one embodiment, the method includes illuminating (by the second endoscope unit) the second field of view with excitation light for the luminescent material that is a fluorescent material, however the excitation light may be of any type (e.g., NIR, visible light, etc.) for any fluorescent material (e.g., exogenous or endogenous, exogenous or endogenous, etc.).

[0117] In one embodiment, the method includes acquiring (with the second endoscope unit) a second sequence of a second set of images including corresponding luminescence images that are representative of luminescence light that is fluorescent light emitted by the luminescent material in the second field of view in response to the excitation light, however, the fluorescent light may be of any type (e.g., NIR, infrared (IR), visible, etc.).

[0118] In one embodiment, the luminescent material is a luminescent agent pre-administered to the patient before carrying out the method. However, the luminescent agent may be of any type (e.g., a targeted luminescent agent based on specific or non-specific interactions, any other non-targeted luminescent agent, etc.), and it may be pre-administered in any manner (e.g., syringe, infusion pump, etc.) at any time (e.g., hours / days before, immediately before carrying out the method, continuously during the method, etc.). Furthermore, the luminescent agent may be administered to the patient in a non-invasive manner (e.g., oral administration for imaging the gastrointestinal tract, by nebulizer to the airway, topical spray application or local introduction during surgery, etc.). In any case, there is no substantial physical intervention of the patient (e.g., intramuscular administration) that requires specialized medical expertise or involves health risks to the patient.

[0119] In one embodiment, the corrections are entered manually according to the display of the first and second images, however, the corrections may be entered according to any display of the first / second images (e.g., the first / second images are played back or paused during acquisition, at the start, and possibly at any subsequent time).

[0120] In one embodiment, the method includes acquiring (with a first endoscope unit) a first sequence of first images that are in color, although the first images may be represented in color in any manner (e.g., any color space, any color model, etc.).

[0121] In one embodiment, the method includes acquiring (with a second endoscopic unit) a second sequence of a second set of images that are in color, although the second images may be rendered in color in any manner (e.g., the same or different from the first images).

[0122] In one embodiment, the method includes preparing the first and second images for said alignment by conversion to grayscale (by a computing device). However, the images may be converted to grayscale in any manner (e.g., based on an average method, a weighting method, etc.).

[0123] In one embodiment, the method includes preparing the first and second images for said registration by downsampling (by a computing device), however, the images may be downsampled at any rate and in any manner (e.g., by simply averaging the values, applying a smoothing filter, etc.).

[0124] In one embodiment, the method includes preparing the first and second images for said alignment by confinement (by a computing device) to corresponding central portions, however, the images may be confined to any central portion (e.g., any size, shape, etc.).

[0125] One embodiment provides a method for training the above neural network, however, this (training) method may be applied to a neural network for any purpose (e.g., its development, maintenance, validation, etc.).

[0126] In one embodiment, the method comprises the following steps under the control of a computer system: However, the computer system may be of any type (see below).

[0127] In one embodiment, the method includes providing (to a computer system) a plurality of first sample images of at least one sample body part of at least one sample patient. However, the first sample images may be any number associated with any number and type of sample body parts of any number and type of sample patients, and may be acquired with any number and type of sample endoscope units inserted into any number of sample cavities of the sample body parts (e.g., the same or different to those described above).

[0128] In one embodiment, the method includes synthesizing (by a computer system) a corresponding composite image from the first sample image, however, the composite image may be synthesized in any manner (e.g., changing shape, changing resolution, changing contrast, zooming in / out, translating, adding / removing noise, any combination thereof, etc.).

[0129] In one embodiment, the method includes synthesizing (by a computer system) a composite image by varying the shape of corresponding first sample images, although the shape may vary in any manner (e.g., from rectangular, polygonal, circular, etc. to circular, rectangular, polygonal, etc.).

[0130] In one embodiment, the method includes synthesizing (by a computer system) a composite image by reducing the resolution of a corresponding first sample image, although the resolution may be reduced to any degree and in any manner (e.g., by downsampling, filtering, etc.).

[0131] In one embodiment, the method includes synthesizing (by a computer system) a composite image by reducing the contrast of corresponding first sample images, although the contrast may be reduced to any degree and in any manner (e.g., by saturating, applying histogram equalization with worsening distributions, etc., at the level of the entire image or a small tile thereof).

[0132] In one embodiment, the method includes synthesizing (by a computer system) a composite image by zooming in on corresponding first sample images, although the first sample images may be zoomed in to any extent and in any manner (e.g., by duplicating values, applying a smoothing filter, etc.).

[0133] In one embodiment, the method includes synthesizing (by a computer system) a composite image by translating corresponding first sample images, although the first sample images may be translated in any range (e.g., by a predefined value, in a random manner, etc.).

[0134] In one embodiment, the method includes synthesizing (by a computer system) the composite image by adding noise to the corresponding first sample image, although the noise may be added in any manner (e.g., adding a predefined noise or random noise uniformly or randomly to the first sample image, etc.).

[0135] In one embodiment, the method includes generating (by a computer system) a plurality of second sample images by each applying a corresponding reference motion to one of the composite images, however, the second sample images may be generated in any number and in any manner (e.g., by using any number and type of reference motions, such as translation, rotation, combinations thereof, by applying all possible reference motions or a corresponding subset thereof selected randomly to each composite image, etc.).

[0136] In one embodiment, the method includes training (by a computer system) a neural network according to a number of training sets each including one of the second sample image, the corresponding first sample image, and the corresponding reference motion. However, the training sets may be selected in any number and in any manner (e.g., random, uniform, etc.), and they may be used to train the neural network in any manner (e.g., based on algorithms such as stochastic gradient descent, real-time iterative learning, high-order gradient descent, extended Kalman filtering, etc.). More generally, the training sets may be provided in any other manner (e.g., by acquiring both the first and second sample images with a corresponding sample endoscope unit, and then realigning each of these pairs manually, automatically, or semi-automatically in other ways, such as those based on the optical flow techniques described above, with the possibility of increasing their number with further pairs of sample images generated by applying random motion, combining both the first and second sample images, applying random motion, etc.).

[0137] In general, similar considerations apply when the same technique is implemented with an equivalent imaging / training method (by using more steps or similar steps with the same functionality in part or by removing some steps that are not required, or by adding further optional steps). Furthermore, steps may be performed (at least partially) in different orders, simultaneously, or in an alternating manner.

[0138] One embodiment provides a computer program which, when executed on a computing device, is configured to cause the computing device to perform a method of operating an endoscopic system to image a body part of a patient. However, the method can be used to image any body part of any patient in any medical procedure (see above). Furthermore, the computer program can be executed on any computing device (see below).

[0139] In one embodiment, the method includes receiving a first sequence of a plurality of first images of a first field of view including at least a portion of the body part. However, the first images may be of any number and type (see above). Furthermore, the first images may be received in any manner (e.g., transferred via a wired / wireless interface, directly acquired, transferred using a removable storage unit, downloaded from a network, etc.).

[0140] In one embodiment, the method includes receiving a second sequence of a plurality of second image sets corresponding to the first image, each of the second image sets including one or more second images of a second field of view including at least a portion of the body part. However, the first images and the second image sets may correspond in any manner, each second image set may include any number of second images of any type, and the first and second fields of view may be of any type (see above). Furthermore, the second images may be received in any manner (e.g., directly acquired, transferred via a wired / wireless interface or using a removable storage unit, downloaded from a network, etc.).

[0141] In one embodiment, the method comprises estimating, for each current image of the plurality of first images and for each current second image of the plurality of second image sets independently, a first motion of a first probe of a first endoscopic unit used to acquire the first sequence of first images and a second motion of a second probe of a second endoscopic unit used to acquire the second sequence of the second image sets, the first motion being estimated according to an estimated set of the plurality of first images corresponding to the current first image and the second motion being estimated according to an estimated set of the plurality of second images corresponding to the current second image. However, these motions may be of any type, be indicated in any way and may be estimated in any way according to an estimated set of first / second images of any type (see above).

[0142] In one embodiment, the method comprises determining at least one correction for each registered pair of a first image of the plurality of first images and a second image of the corresponding second image set according to a first movement and a second movement of one or more calculated sets of the plurality of first images and the plurality of second images, respectively, corresponding to the pair of first and second images. However, the correction may be of any type and it may be determined in any manner (see above).

[0143] In one embodiment, the method includes a step of registering each of the first and second images according to a corresponding correction, however, the registration may be of any type and may be performed in any manner and at any time (see above).

[0144] In one embodiment, the method includes outputting a representation of the body part based on the aligned first and second images of each aligned pair, however, the representation of the body part may be output in any manner, to any output unit, and based on images aligned in any manner (see above).

[0145] An embodiment provides a computer program product comprising a computer readable storage medium embodying a computer program, the computer program being loadable into a working memory of a computing device, thereby configuring the computing device to perform the same (operating) method when the computer program is executed on the computing device.

[0146] One embodiment provides a computer program arranged to cause a computer system to carry out the above (training) method when the computer program is executed on the computer system, although the computer program may be executed on any computer system (see below).

[0147] One embodiment provides a computer program product comprising a computer readable storage medium embodying a computer program, the computer program being loadable into a working memory of a computer system, thereby configuring the computer system to perform the same (training) method when the computer program is executed on the computer system.

[0148] In general, each (computer) program may be implemented as a stand-alone module, as a plug-in, or in the latter case directly, to an existing software program (e.g., an imaging application of a computer device for an operating method, or a configuration application of a computer system for a training method). In either case, similar considerations apply if the program is configured in a different way, or if additional modules or functions are provided. Similarly, the memory structures may be of other types, or may be replaced by equivalent entities (not necessarily consisting of physical storage media). The programs may take any form suitable for use by any computer device / system (see below), thereby configuring the computer device / system to perform the desired operations. In particular, the programs may be in the form of external or resident software, firmware or microcode (either object code or source code, e.g., compiled or interpreted). Furthermore, it is possible to provide a computer program on any computer-readable storage medium. The storage medium is any tangible medium (as opposed to a transitory signal itself) that can hold and store instructions used by the computer device / system. For example, the storage medium may be of electronic, magnetic, optical, electromagnetic, infrared or semiconductor type. Examples of such storage media are fixed disks (which may be preloaded with the program), removable disks, memory keys (e.g., USB type), etc. The program may be downloaded to the computing device / system from the storage medium or over a network (e.g., the Internet, a wide area network, and / or a local area network, including transmission cables, optical fibers, wireless connections, network devices). One or more network adapters in the computing device / system receive the program from the network and transfer it to one or more storage devices of the computing device / system for storage.In any case, the techniques according to an embodiment of the present disclosure lend themselves to implementation in a hardware structure (e.g., by electronic circuitry integrated in one or more chips of semiconductor material, such as a field programmable gate array (FPGA) or application specific integrated circuit (ASIC)), or in combination with suitably programmed or otherwise configured software and hardware.

[0149] One embodiment provides an endoscopic system for performing the steps of the above (imaging) method, however the endoscopic system may be of any type (e.g. endoscopic units connected together, one of which comprises a computing device and an output unit, endoscopic units connected to a separate computing device and output unit, all components integrated together, etc.).

[0150] In one embodiment, the endoscopic system comprises a first endoscope unit for acquiring a first sequence of first images, however the first endoscope unit may be of any type (e.g., a motherscope, a separate endoscope based on any number and types of lenses, waveguides, mirrors, sensors, etc.).

[0151] In one embodiment, the endoscopic system comprises a second endoscope unit for acquiring a second sequence of a second image set, however, the second endoscope unit may be of any type (e.g., a baby scope, a separate endoscope based on any number and types of lenses, waveguides, mirrors, sensors, etc.).

[0152] In one embodiment, the endoscopic system comprises a computing device for registering each registered pair of first and second images, however the computing device may be of any type (see below).

[0153] In one embodiment, the endoscopy system comprises an output unit for outputting a representation of the body part, however the output unit may be of any type (e.g. a monitor, virtual reality glasses, a printer, etc.).

[0154] One embodiment provides an endoscope device (for use in the endoscope system described above), which comprises one between the first endoscope unit and the second endoscope unit, however the endoscope device may be of any type (e.g. comprising a baby scope, a mother scope, etc.).

[0155] In one embodiment, the endoscopic unit comprises an acquisition unit for acquiring a corresponding one between a first sequence of a first image set and a second sequence of a second image set, however the acquisition unit may be of any type (e.g. based on CCD, ICCD, EMCCD, CMOS, InGaAs or PMT sensors, any illumination unit for applying excitation light to the body part, e.g. based on laser, LED, UV lamp etc. and / or white light, e.g. LED, halogen / xenon lamp etc.).

[0156] In one embodiment, the endoscopic unit comprises an interface for receiving the other of the first sequence of the first image set and the second sequence of the second image set from the other of the first endoscopic unit and the second endoscopic unit, although the interface may be of any type (e.g., wired, wireless, serial, parallel, etc.).

[0157] In one embodiment, the endoscopic unit comprises a computing device for registering each registered pair of first and second images, however the computing device may be of any type (see below).

[0158] In one embodiment, the endoscopy unit comprises an output unit for outputting a representation of the body part, however the output unit may be of any type (see above).

[0159] An embodiment provides a computing device, which comprises means configured to perform the steps of the above (operational) method. An embodiment provides a computing device comprising circuitry (i.e. any hardware suitably configured, e.g. by software) for performing each step of the same (operational) method. However, the computing device may be of any type (e.g. a central unit of each endoscopic unit receiving the other image from the other endoscopic unit, a common central unit of the endoscopic system for both endoscopic units, a separate computer receiving corresponding images from the two endoscopic units or receiving all images from the endoscopic system, etc.).

[0160] An embodiment provides a computer system comprising means configured to perform the steps of the above (training) method. An embodiment provides a computer system including circuitry (i.e. any hardware suitably configured, for example, by software) for performing each step of the same (training) method. However, the computer system may be of any type (for example, a personal computer, a server, a virtual machine provided, for example, in a cloud environment, etc.).

[0161] In general, similar considerations apply when the endoscopic system, endoscopic equipment, computer device, and computer system each have different structures or have equivalent components, or when they have other operational characteristics. In either case, the entire components may be separated into more elements, or two or more components may be integrated into a single element. Furthermore, each component may be replicated to support the execution of corresponding operations in parallel. Furthermore, unless otherwise specified, any interaction between different components generally does not have to be sequential, but may be direct or indirect through one or more intermediaries.

[0162] An embodiment provides a surgical method comprising the following steps: a body part of a patient is imaged by executing said (imaging) method, thereby outputting a representation of the body part during a surgical treatment of the body part; the body part is operated on according to the output of said representation; however, the proposed method may find application in any kind of surgical method in the broadest sense of the term (e.g. for therapeutic purposes, for preventive purposes, for aesthetic purposes, etc.) and for acting on any kind of body part of any patient (see above).

[0163] An embodiment provides a diagnostic method comprising the following steps: A body part of a patient is imaged by executing the (imaging) method described above, thereby outputting a representation of the body part during a diagnostic procedure of the body part; A health state of the body part is evaluated according to the output of said representation. However, the proposed method may find application in any kind of diagnostic method in the broadest sense of the term (e.g. for discovering new lesions, for monitoring known lesions, etc.) and for analysing any kind of body part of any patient (see above).

[0164] An embodiment provides a therapeutic method comprising the following steps: a body part of a patient is imaged by executing said (imaging) method, thereby outputting a representation of the body part during a therapeutic treatment of the body part; the body part is treated according to the output of said representation; however, the proposed method may find application in any kind of therapeutic method in the broadest sense of the term (e.g. for treating a pathological condition, avoiding its progression, preventing the occurrence of a pathological condition or simply improving the comfort of the patient) and for acting on any kind of body part of any patient (see above).

Claims

1. A method (500) for imaging a body part (103) of a patient (106) using an endoscope system (100), comprising: acquiring (515) a first sequence of a plurality of first images of a first field of view (221m) including at least a portion of the body part (103) using a first endoscope unit (115m) of the endoscope system (100) via a first probe (127m) of the first endoscope unit (115m); acquiring (509-512) a second sequence of a plurality of second image sets corresponding to the first image using a second endoscope unit (115b) of the endoscope system (100) via a second probe (127b) of the second endoscope unit (115b) that is movable relative to the first probe (127m), each of the second image sets including one or more second images of a second field of view (221b) that includes at least a portion of the body part (103); registering (518-584) each registered pair of a first image of the plurality of first images and a corresponding second image of the second set of images by the computing device (242b); and outputting (587) to an output unit (121b) a representation of the body part (103) based on the first and second images of each registered pair.

2. The method (500) of claim 1, further comprising steps (509-512) of acquiring a second sequence of a second image set using a second endoscope unit (115b) via a second probe (127b) removably inserted into the working channel (112) of the first endoscope unit (115m).

3. the computer device (242b) is included in one (115b) between the first endoscope unit (115m) and the second endoscope unit (115b); The method (500) according to claim 1 or 2, comprising a step (115) of receiving, by the computing device (242b), from the other (115m) of the first endoscope unit (115m) and the second endoscope unit (115b), a corresponding first sequence of a first image set or a second sequence of a second image set.

4. The method (500) includes determining, by the computing device (242b), at least one correction (539-542; 554-563; 578-581) for each registered pair of first and second images due to corresponding relative motion of the first and second probes (127m) and (127b); and registering (548; 569; 584) each registered pair of first and second images according to the corresponding correction by a computing device (242b).

5. The method (500) includes estimating (554-557; 539-542) by the computing device (242b) a first motion of the first probe (115m) and a second motion of the second probe (115b) for each current image of the plurality of first images and for each current second image of the plurality of second images, respectively, independently, wherein the first motion is estimated according to an estimated set of the plurality of first images corresponding to the current first image, and the second motion is estimated according to an estimated set of the plurality of second images corresponding to the current second image; and determining, by the computing device (242b), corrections for each registered pair of first and second images according to first and second movements of one or more calculated sets of the plurality of first images and the plurality of second images, respectively, corresponding to the pair of first and second images.

6. The method (500) of claim 5, comprising steps (554-557) of estimating, by the computing device (242b), a first motion vector indicative of a first motion for each current first image according to the estimated set of corresponding first images, and a second motion vector indicative of a second motion for each current second image according to the estimated set of corresponding second images.

7. The method (500) includes calculating (560), by the computing device (242b), a misalignment of each pair of corresponding current first and second images according to the corresponding first and second motions; The method (500) of claim 5, further comprising: calculating (563) by the computing device (242b) a correction for each aligned pair of first and second images from the aligned pair of first and second images according to the misalignment of the corresponding calculated set of first and second images, weighted in a decreasing manner with the corresponding range and / or the corresponding temporal distance.

8. The method (500) includes providing (539), by the computing device (242b), input data to the neural network based on the further set of estimates of the first image and the further set of estimates of the second image for each registered pair of the first and second images; and receiving, by the computing device, a correction for each registered pair of the first and second images from the neural network.

9. 9. The method (500) of claim 8, further comprising: selecting (542) a correction for each registered pair of the first and second images from among a plurality of predefined corrections by a neural network (439).

10. The neural network (439) is a convolutional neural network, 10. The method of claim 8, wherein the convolutional neural network includes a plurality of groups along a processing direction of the convolutional neural network, each group including one or more convolutional layers followed by a max-pooling layer and a plurality of fully-connected layers providing corresponding probabilities of predefined corrections.

11. The method (500) includes the steps of: acquiring (515) with a first endoscope unit (115m) a first sequence of first images, each of the first images including a plurality of first values ​​representing a corresponding location of a first field of view (221m); acquiring (509-512) a second sequence of second image sets using a second endoscope unit (115b), each image set including a plurality of second values ​​representing a corresponding location of a second field of view (221b); 9. The method (500) of claim 8, further comprising: providing (539), by the computing device (242b), input data for each registration pair of the first and second images to the neural network, the input data being obtained by concatenating, for each location, a first value of the first image of the further estimation set and a second value of the second image of the further estimation set.

12. 5. The method (500) of claim 4, further comprising receiving (578-581) by a computing device (242b) a correction manually input in accordance with a representation of the first image and the second image.

13. The method (500) includes the steps of acquiring (515) with a first endoscope unit (115m) a first sequence of first images, the first sequence being corresponding reflectance images representative of visible light reflected by content in a first field of view (221m); acquiring (509-512) a second sequence of second image sets using the second endoscope unit (115b), the second image sets being corresponding luminescence images representing luminescence light emitted by luminescent materials in the second field of view (221b) and corresponding further reflectance images representing visible light reflected by the contents of the second field of view (221b); determining, by the computing device (242b), a correction for each registered pair of reflectance and luminescence images according to the corresponding reflectance image and the further reflectance image (539-542; 554-563); registering (548; 569) each registered pair of reflectance and luminescence images according to the corresponding corrections by the computing device (242b); 5. The method (500) of claim 4, further comprising: outputting (587) to an output unit (121b) a representation of the body part (103) based on the reflectance image and luminescence image of each registered pair.

14. The method (500) includes the steps of acquiring (515) with a first endoscope unit (115m) a first sequence of first images, the first sequence being corresponding reflectance images representative of visible light reflected by content in a first field of view (221m); The method (500) of claim 1 further comprises a step (509) of acquiring, using a second endoscope unit (115b), a second sequence of a second set of images, which are corresponding luminescent images representing luminescent light emitted by the luminescent material within a second field of view (221b).

15. The method (500) includes the step of illuminating (506) the second field of view (221b) with excitation light for the luminescent material, which is a fluorescent material, by the second endoscope unit (115b); The method (500) of claim 13 further comprises a step (509) of acquiring, using a second endoscope unit (115b), a second sequence of a second image set including corresponding luminescence images which are corresponding fluorescence images representing luminescence light which is fluorescent light emitted by a fluorescent substance within a second field of view in response to the excitation light.

16. 14. The method (500) of claim 13, wherein the luminescent substance is a luminescent agent that is pre-administered to the patient (106) prior to performing the method (500).

17. The method (500) includes the steps of acquiring (515) a first sequence of first images, which are in color, using a first endoscope unit (115m); acquiring (509-512) a second sequence of second images, which are in color, using a second endoscope unit (115b); and preparing (533), by a computing device (242b), the first and second images for said alignment (536-572) by conversion to grayscale, downsampling, and / or limiting to corresponding central portions.

18. 10. A method (800) for training a neural network (439) for use in a method (500) according to any of claims 8, comprising: The method (800) comprises, under the control of a computer system (600): Providing (810) to a computer system (600) a plurality of first sample images of at least one sample body part (103) of at least one sample patient (106); synthesizing (815), by the computer system (600), a corresponding composite image from the first sample image by changing the shape of the corresponding first sample image, reducing its resolution, reducing its contrast, zooming in on it, translating it, and / or adding noise to it; generating (820) a plurality of second sample images by applying, by the computer system (600), a corresponding reference motion to one of the composite images respectively; and training (825-880) the neural network (439) by the computer system (600) according to a plurality of training sets each including one of the second sample image, the corresponding first sample image, and the corresponding reference motion.

19. 1. A computer program (400) configured, when executed on a computing device (242b), to cause the computing device (242b) to perform a method (500) for operating an endoscopy system (100) to image a body part (103) of a patient (106), the method (500) comprising: receiving (515) a first sequence of a plurality of first images of a first field of view (221m) including at least a portion of the body part (103); receiving (509-512) a second sequence of a plurality of second image sets corresponding to the first image, each of the second image sets including one or more second images of a second field of view (221b) including at least a portion of the body part (103); and estimating (554-557; 539-542) for each current image of the plurality of first images and for each current second image of the plurality of second image sets, independently, a first motion of a first probe (115m) of a first endoscope unit (115m) used to acquire the first sequence of first images and a second motion of a second probe (115b) of a second endoscope unit (115b) used to acquire the second sequence of second images, wherein the first motion is estimated according to the estimated set of the plurality of first images corresponding to the current first image and the second motion is estimated according to the estimated set of the plurality of second images corresponding to the current second image; determining (560-563; 539-542) at least one correction for each registered pair of a first image of the plurality of first images and a second image of the corresponding second set of images according to a first motion and a second motion of one or more calculated sets of the plurality of first images and the plurality of second images, respectively, corresponding to the pair of first and second images; - registering each registered pair of first and second images according to the corresponding correction (548; 569); and outputting (587) a representation of the body part (103) based on the first image and the second image of each registered pair.

20. 1. A computer program product comprising a computer-readable storage medium embodying a computer program, a computer program loadable into a working memory of a computing device such that, when the computer program is executed on the computing device, the computing device is configured to perform a method of operating an endoscopy system to image a body part of a patient; The method includes receiving a first sequence of a plurality of first images of a first field of view that includes at least a portion of a body part; receiving a second sequence of a plurality of second image sets corresponding to the first image, each of the second image sets including one or more second images of a second field of view including at least a portion of the body part; estimating, for each current image of the plurality of first images and for each current second image of the plurality of second image sets, independently, a first motion of a first probe of a first endoscope unit used to acquire the first sequence of first images and a second motion of a second probe of a second endoscope unit used to acquire the second sequence of second images, wherein the first motion is estimated according to the estimated set of the plurality of first images corresponding to the current first image and the second motion is estimated according to the estimated set of the plurality of second images corresponding to the current second image; determining at least one correction for each registered pair of a first image of the plurality of first images and a second image of the corresponding second set of images according to first and second motions of one or more calculated sets of the plurality of first images and the plurality of second images, respectively, corresponding to the pair of first and second images; registering each registered pair of first and second images according to the corresponding correction; and outputting a representation of the body part based on the first image and the second image of each registered pair.

21. A computer program (700) configured to cause a computer system (600) to perform the method (800) of claim 18 when the computer program is executed on the computer system (600).

22. 1. A computer program product comprising a computer-readable storage medium embodying a computer program, A computer program product, the computer program being loadable into a working memory of a computer system, whereby when the computer program is executed on the computer system, it configures the computer system to carry out the method of claim 18.

23. An endoscope system (100) for performing the steps of the method (500) of claim 1, comprising: a first endoscope unit (115m) for acquiring a first sequence of first images; a second endoscope unit (115b) for acquiring a second sequence of a second set of images; a computing device (242b) for registering each registered pair of the first and second images; and an output unit (121b) for outputting a representation of the body part (103).

24. An endoscopic device (115b) for use in the endoscopic system according to claim 23, comprising: The endoscope device (115b) includes one of a first endoscope unit (115m) and a second endoscope unit (115b), The endoscope unit (115b) an acquisition unit (224b-239b) for acquiring a corresponding one between a first sequence of first images and a second sequence of a second set of images; an interface (124b) for receiving the first sequence of first images and the second sequence of second images from the other (115m) of the first endoscope unit (115m) and the second endoscope unit (115b); a computing device (242b) for registering each registered pair of the first and second images; an output unit (121b) for outputting a representation of the body part (103) based on the first image and the second image of each registered pair.

25. a computing device (242b) for operating the endoscopy system (100) to image a body part (103) of a patient (106), the computing device (242b) comprising: means (412) for receiving (515) a first sequence of a plurality of first images of a first field of view (221m) including at least a portion of the body part (103); means (403) for receiving (509-512) a second sequence of a plurality of second image sets corresponding to the first images, each second image set including one or more second images of a second field of view (221b) including at least a portion of the body part (103); means (427; 439) for estimating (554-557; 539-542), for each current image of the plurality of first images and for each current second image of the plurality of second image sets, independently, a first motion of a first probe (115m) of a first endoscope unit (115m) used to acquire the first sequence of first images and a second motion of a second probe (115b) of a second endoscope unit (115b) used to acquire the second sequence of second images, wherein the first motion is estimated according to the estimated set of the plurality of first images corresponding to the current first image and the second motion is estimated according to the estimated set of the plurality of second images corresponding to the current second image; means (426; 439) for determining (560-563; 539-542) at least one correction for each registered pair of a first image of the plurality of first images and a second image of a corresponding second set of images according to a first motion and a second motion of one or more calculated sets of the plurality of first images and the plurality of second images, respectively, corresponding to the pair of first and second images; means (448) for registering (548; 569) each registered pair of first and second images according to the corresponding correction; and means (454) for outputting (587) a representation of the body part (103) based on the first image and the second image of each registered pair.

26. 1. A computing device for operating an endoscopy system to image a body part of a patient, the computing device comprising: a circuit for receiving a first sequence of a plurality of first images of a first field of view including at least a portion of the body part; a circuit for receiving a second sequence of a plurality of second image sets corresponding to the first image, each of the second image sets including one or more second images of a second field of view (221b) including at least a portion of the body part (103); a circuit for estimating, independently for each current image of the plurality of first images and for each current second image of the plurality of second image sets, a first motion of a first probe of a first endoscope unit used to acquire the first sequence of first images and a second motion of a second probe of a second endoscope unit used to acquire the second sequence of second image sets, wherein the first motion is estimated according to the estimated set of the plurality of first images corresponding to the current first image and the second motion is estimated according to the estimated set of the plurality of second images corresponding to the current second image; a circuit for determining at least one correction for each registered pair of a first image of the plurality of first images and a second image of a corresponding second set of images according to a first motion and a second motion of one or more calculated sets of the plurality of first images and the plurality of second images, respectively, corresponding to the pair of first and second images; a circuit for aligning each registered pair of a first image and a second image according to the corresponding correction; and circuitry for outputting (587) a representation of the body part (103) based on the first and second images of each registered pair.

27. A computer system (600) comprising means (700) configured to perform the steps of the method (800) of claim 18.

28. A computer system (600) comprising circuitry configured to perform the steps of the method of claim 18.