Image Processing Method and Apparatus

Through the combination of UNet deep learning network and TapNet-Siamese key point tracking network, the problem of short time window of ICG fluorescence imaging technology is solved, and the accuracy and stability of boundary extraction in a dynamic environment is achieved, and the boundary prompt time is extended.

CN120107301BActive Publication Date: 2025-07-15CHONGQING FUDIMAI DIGITAL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510584722.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-15
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The short observation time window of ICG fluorescence imaging technology makes it difficult to accurately calibrate organ or tumor boundaries in a very short time during laparoscopic surgery. The prior art lacks robustness and accuracy under dynamic laparoscopic field of vision and complex tissue structures.

Method used

The trained UNet deep learning network is used for global segmentation, combined with Canny edge detection and morphological operations, the probability map of the basin boundary is obtained, and the network reconstruction boundary is tracked through TapNet-Siamese key points to achieve stable prompts for dynamic changes.

Benefits of technology

It improves the accuracy and stability of boundary extraction, extends the boundary prompt time, can effectively deal with organ turn and deformation, and ensures the accuracy of the surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107301B_ABST
    Figure CN120107301B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image processing method and apparatus, belonging to the technical field of data processing. The method includes: acquiring an ICG fluorescence imaging image; globally segmenting the ICG fluorescence imaging image by using a trained UNet deep learning network to obtain a probability map of the watershed boundary; processing the probability map to obtain a final boundary; sampling the final boundary to obtain a set of key point coordinates; determining a displacement vector of the set of key point coordinates within a preset time interval; and reconstructing the dynamic change of the final boundary according to the set of key point coordinates and the displacement vector. The present disclosure realizes the delayed display of the fluorescence boundary and has excellent dynamic adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of image processing, and particularly relates to an image processing method and apparatus. Background Art

[0002] With the development of minimally invasive surgery technology, laparoscopic surgery, as a precise and minimally invasive treatment method, has been increasingly widely used in clinical applications. During laparoscopic surgery, how to accurately calibrate the boundaries of organs, tumors, or other anatomical structures plays a crucial role in improving the safety and treatment effect of the surgery.

[0003] Fluorescence imaging technology, especially using ICG (indocyanine green) fluorescent dye for boundary calibration, has become an important means for intraoperative boundary recognition.

[0004] ICG fluorescence imaging technology has good imaging effects, but its optimal observation time window is short. The development time of ICG is usually 35 seconds, between about 29 seconds and 44 seconds, and its development duration is about 3 minutes. The diffusion of ICG in the microcirculation will cause the watershed boundary to gradually become blurred. Therefore, the optimal observation window is usually within 30 seconds to 1 minute after staining. The shortness of the time window poses certain challenges to precise operations during the surgery. Especially in complex laparoscopic surgeries, it is necessary to complete the marking of the watershed boundary within an extremely short time to ensure the accuracy of the surgery.

[0005] The Chinese patent document with the publication number CN112414987A discloses a technology with the name "A fluorescence imaging device and test method applying spectral detection". This technology cannot accurately extract boundary details within the fluorescence imaging window, and it is difficult to solve the situation of boundary blurring or disappearance. Existing boundary detection technologies mostly rely on traditional image processing methods and lack sufficient robustness and accuracy when facing dynamic laparoscopic fields and complex tissue structures. The Chinese patent document with the publication number CN101770141A discloses a technology with the name "A fluorescence navigation system during tumor surgery". Although this technology can achieve boundary tracking to a certain extent, it has poor adaptability to dynamic changes such as organ flipping and deformation, and cannot ensure the continuous stability of boundary prompts. Summary of the Invention

[0006] The present disclosure provides an image processing method and apparatus, aiming to solve the problem of the short observation time window in ICG fluorescence imaging technology.

[0007] According to a first aspect of the present disclosure, there is provided an image processing method, the method comprising: acquiring an ICG fluorescence imaging image; globally segmenting the ICG fluorescence imaging image by using a trained UNet deep learning network to obtain a probability map of the watershed boundary; processing the probability map to obtain a final boundary; sampling the final boundary to obtain a set of key point coordinates , specifically including: selecting no less than a preset number of anchor points at the curvature extreme points according to a confidence threshold , and satisfying that the adjacent anchor points are less than or equal to to form the set of key point coordinates , where the local curvature corresponding to the anchor point is the curvature extreme value of the corresponding pixel on the curve, and the confidence of the final boundary is greater than or equal to , , and the arc length difference between adjacent anchor points along the curve is less than or equal to ; determining the set of key point coordinates of the displacement vector within a preset time interval by using an improved TapNet-Siamese key point tracking network, that is , specifically including: performing feature embedding on local image patches centered on each key point in the reference frame and the current frame through a double-branch convolutional subnetwork with shared weights; obtaining the displacement vector by using cosine similarity to match the embedding vectors; when the matching confidence is lower than the threshold or the cross-frame displacement is greater than , calling a cross-frame cross-attention network based on the TapNet structure to re-estimate the displacement vector, where represents the displacement threshold, and represents the confidence threshold; the TapNet-Siamese key point tracking network includes: an input layer for receiving the reference frame, the current frame, and the key point coordinates, a double-branch convolutional-TSM-ResNet18 backbone with completely shared weights, a Siamese metric learning branch for outputting key point embedding vectors, a Cross-Attention module for performing cross-frame feature alignment when the confidence is lower than the threshold , and a displacement regression head for outputting ; if the average confidence of consecutive K frames is lower than the threshold , recalculating the probability map and refreshing the set of key point coordinates ; reconstructing the dynamic change of the final boundary according to the set of key point coordinates and the displacement vector .

[0008] In some embodiments, processing the probability map to obtain the final boundary includes: performing binarization on the probability map to obtain a rough boundary; using the Canny edge detection algorithm to detect the ICG fluorescence imaging image to obtain a detailed boundary; fusing the rough boundary and the detailed boundary to obtain a fused boundary; and processing the fused boundary by means of a morphological operation function to obtain the final boundary.

[0009] In some embodiments, before obtaining the ICG fluorescence imaging image, it further includes: collecting an initial ICG fluorescence imaging image; performing grayscale processing on the initial ICG fluorescence imaging image to obtain a grayscale image; removing noise from the grayscale image to obtain the ICG fluorescence imaging image; and returning the ICG fluorescence imaging image.

[0010] In some embodiments, using the trained UNet deep learning network to globally segment the ICG fluorescence imaging image to obtain a probability map of the watershed boundary, wherein the cross-entropy loss function is used to optimize the UNet deep learning network, and the Adam optimizer is used to update the parameters of the UNet deep learning network.

[0011] In some embodiments, using the trained UNet deep learning network to globally segment the ICG fluorescence imaging image to obtain a probability map of the watershed boundary includes: according to the formula: , obtaining the probability map , where represents the trained UNet deep learning network, represents the ICG fluorescence imaging image, represents the spatial coordinates, represents the time variable.

[0012] In some embodiments, performing binarization on the probability map to obtain a rough boundary includes: according to the formula: , obtaining the rough boundary , where represents the set threshold.

[0013] In some embodiments, using the Canny edge detection algorithm to detect the ICG fluorescence imaging image to obtain a detailed boundary includes: according to the formula: , obtaining the detailed boundary .

[0014] In some embodiments, processing the fused boundary by means of a morphological operation function to obtain the final boundary includes: according to the formula: , obtaining the final boundary , where Indicates the fusion boundary, Indicates the morphological operation function.

[0015] In some embodiments, reconstructing the dynamic change of the final boundary according to the set of key point coordinates and the displacement vector includes: According to the formula: , reconstructing the dynamic change of the final boundary, where, Indicates a preset time interval The displacement vectors of each key point within Indicates at the moment The set of key point coordinates sampled, Indicates the set of key point coordinates after dynamic change .

[0016] According to a second aspect of the present disclosure, there is provided an image processing apparatus, including: a memory; and a processor coupled to the memory, the processor being configured to execute the image processing method as described above based on instructions stored in the memory.

[0017] By adopting the above technical solutions, the beneficial technical effects that can be achieved by the embodiments of the present disclosure are as follows: The trained UNet deep learning network is used to globally segment the ICG fluorescence imaging image to obtain the probability map of the watershed boundary. The Canny edge detection algorithm is used to detect the ICG fluorescence imaging image to obtain the detailed boundary. The morphological operation function is used to process the fusion boundary to obtain the final boundary, improving the accuracy of boundary extraction in a dynamic and low signal-to-noise ratio environment. According to the set of key point coordinates and the displacement vector, the dynamic change of the final boundary is reconstructed, which can effectively cope with the organ flipping and deformation situations, and maintain the continuous stability of the boundary prompt, that is, extend the boundary prompt time.

[0018] The present disclosure adds a key point screening rule: N anchor points are selected according to the local maximum curvature of the final boundary plus the confidence threshold, and N is greater than or equal to 50, and the arc length difference between adjacent anchor points is less than or equal to , traditional optical flow does not involve adaptive sampling of segmentation confidence, and samples are encrypted in the complex curve area, improving the reconstruction accuracy and reducing the redundant point bandwidth.

[0019] Based on the original TapNet framework, this disclosure introduces an explicit Siamese (weight-sharing) metric learning branch, enhancing the long-term key point re-identification and occlusion recovery capabilities. The original TapNet focuses on cross-frame feature fusion, and the backbone network structure is a dual-frame joint encoder (TSM-ResNet18 + Cross-Attention). Regarding two problems in the intraoperative scenario, namely, occlusion / great displacement after fluorescence attenuation: simple cross-attention is prone to mismatch, and frame rate plus resolution limitation: the original TapNet has a large full-image-level interaction calculation amount, this disclosure proposes a TapNet-Siamese key point tracking network, which supplements local Patch descriptors for each key point; if the cross-frame is greater than or the matching confidence is less than , then Siamese Re-ID is enabled; the Siamese branch performs lightweight embedding comparison on the Patch; only when the confidence is low, it returns to the backbone for re-segmentation. It is stipulated to use the end-to-end network of Siamese TapNet, output and with a re-identification and re-localization mechanism (automatically return to UNet re-segmentation when the score is less than ). Compared with optical flow (local gradient method) and TapNet (global metric learning), this disclosure provides a tracking retention rate of greater than or equal to 95% under cross-occlusion / disparity / light drift.

[0020] In this disclosure, if the average anchor confidence is less than for continuous K frames, then re-segmentation is automatically triggered, introducing a closed-loop self-recovery mechanism, and the delay display window is extended from 3 minutes to more than 8 minutes. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0022] With reference to the drawings, the present disclosure can be more clearly understood according to the following detailed description.

[0023] Figure 1 is a flowchart showing an image processing method according to some embodiments of the present disclosure.

[0024] Figure 2 is a flowchart showing boundary extraction according to some embodiments of the present disclosure.

[0025] Figure 3 is a schematic diagram showing the reconstruction of a boundary with dynamic changes according to some embodiments of the present disclosure.

[0026] Figure 4 is a schematic diagram showing an image display system according to some embodiments of the present disclosure.

[0027] Figure 5 It is a schematic diagram showing the time-lapse display effect of a fluorescence imaging image according to some embodiments of the present disclosure.

[0028] Figure 6 It is a block diagram showing an image processing apparatus according to some embodiments of the present disclosure.

[0029] Figure 7 It is a block diagram showing an image processing apparatus according to some other embodiments of the present disclosure.

[0030] Figure 8 It is a block diagram showing a computer system for implementing some embodiments of the present disclosure. Detailed Embodiments

[0031] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0032] Meanwhile, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0033] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or its uses.

[0034] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.

[0035] In all the examples shown and discussed here, any specific values should be understood as merely exemplary and not as limitations. Therefore, other examples of the exemplary embodiments may have different values.

[0036] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0037] Currently, with the development of minimally invasive surgical techniques, laparoscopic surgery, as a precise and minimally invasive treatment method, is becoming increasingly widespread in clinical applications. During laparoscopic surgery, how to accurately calibrate the boundaries of organs, tumors, or other anatomical structures plays a crucial role in improving the safety and treatment effect of the surgery.

[0038] Fluorescence imaging technology, especially using ICG (indocyanine green) fluorescent dye for boundary calibration, has become an important means for intraoperative boundary recognition.

[0039] The ICG fluorescence imaging technology has good imaging effects, but its optimal observation time window is short. The development time of ICG is usually 35 seconds, between approximately 29 seconds and 44 seconds, and its development duration is about 3 minutes. The diffusion of ICG in the microcirculation causes the basin boundaries to gradually blur. Therefore, the optimal observation window is usually within 30 seconds to 1 minute after staining. The shortness of the time window poses a certain challenge to precise operations during the surgery. Especially in complex laparoscopic surgeries, it is necessary to complete the marking of the basin boundaries within an extremely short time to ensure the accuracy of the surgery.

[0040] In view of this, the present disclosure proposes an image processing method and apparatus. The trained UNet deep learning network is used to globally segment the ICG fluorescence imaging image to obtain the probability map of the basin boundaries. The Canny edge detection algorithm is used to detect the ICG fluorescence imaging image to obtain the detailed boundaries. The morphological operation function is used to process the fused boundaries to obtain the final boundaries, improving the accuracy of boundary extraction in dynamic and low signal-to-noise ratio environments. According to the key point coordinate set and displacement vectors, the dynamic changes of the final boundaries are reconstructed, which can effectively cope with the flipping and deformation of organs and maintain the continuous stability of the boundary prompts, that is, extend the boundary prompt time.

[0041] Figure 1 It is a flowchart showing the image processing method according to some embodiments of the present disclosure. As Figure 1 shown, the image processing method includes steps S110 to S170.

[0042] In step S110, an ICG fluorescence imaging image is acquired.

[0043] The ICG fluorescence imaging technology is a near-infrared optical imaging method based on indocyanine green dye, which provides precise navigation for surgeries by real-time displaying the fluorescence signals of tissues or lesions.

[0044] Before acquiring the ICG fluorescence imaging image, it further includes: collecting the initial ICG fluorescence imaging image; performing grayscale processing on the initial ICG fluorescence imaging image to obtain a grayscale image; removing noise from the grayscale image to obtain the ICG fluorescence imaging image; and returning the ICG fluorescence imaging image.

[0045] As Figure 2 shown, according to the formula: , grayscale processing is performed, where represents grayscale processing.

[0046] According to the formula: , noise removal is performed, where represents noise removal.

[0047] Image grayscale processing is the process of converting a color image into a grayscale image, that is, by mathematical methods, the red, green, and blue color channel values of each pixel are converted into a single grayscale value. The grayscale value range of each pixel in a grayscale image is 0 - 255, where 0 is pure black and 255 is pure white. By eliminating color information and only retaining luminance information, the core formula for grayscale conversion can be expressed as: , where represents the conversion function, and different methods correspond to different calculation logics. The purpose of grayscale processing is to simplify data, enhance contrast, and improve efficiency. That is, a grayscale image only requires single-channel storage, and the data volume is reduced to 1 / 3 of the original image. Color interference is removed to highlight the image structure and enhance contrast. The computational complexity of subsequent processing, such as edge detection and image segmentation, is reduced.

[0048] The main purpose of denoising is to restore the true information of the image by eliminating or suppressing noise interference and improve the reliability of subsequent image processing and analysis. For example, eliminating high-frequency noise interference enables the Canny algorithm to more accurately locate edges. For example, suppressing artifacts caused by noise improves the robustness of the segmentation algorithm.

[0049] In step S120, the trained UNet deep learning network is used to globally segment the ICG fluorescence imaging image to obtain a probability map of the watershed boundary.

[0050] The global segmentation of the ICG fluorescence imaging image by the UNet deep learning network aims to solve the key clinical needs in medical image segmentation through high-precision and real-time pixel-level classification. Taking tumor boundary recognition as an example, ICG retains fluorescence in the tumor area due to abnormal metabolism. UNet extracts multi-scale features through an encoder-decoder structure and combines skip connections to fuse shallow details and deep semantic information, enabling accurate segmentation of the tumor boundary, especially performing outstandingly in low signal-to-noise ratio scenarios.

[0051] Using the trained UNet deep learning network to globally segment the ICG fluorescence imaging image to obtain a probability map of the watershed boundary, where the cross-entropy loss function is used to optimize the UNet deep learning network, and the Adam optimizer is used to update the parameters of the UNet deep learning network.

[0052] The main functions of the cross-entropy function in UNet include pixel-level classification optimization, handling class imbalance, gradient characteristics, and model convergence.

[0053] The core goal of UNet is to perform pixel-level segmentation on images. Cross-entropy drives the update of network parameters by quantifying the difference between the predicted probability distribution and the true label distribution of each pixel, and its mathematical expression is , where represents the true label, Represents the class probabilities predicted by the model. This loss function is particularly suitable for binary classification (such as lesion / normal tissue) or multi-classification scenarios.

[0054] In medical images, the proportion of the lesion area is usually small. Cross-entropy alleviates class imbalance through the following strategies: First, weighted cross-entropy: assigns higher weights to the minority classes. For example, in tumor segmentation, it enhances the loss contribution of lesion pixels. Second, focal loss: reduces the weights of easily classified samples by adjusting the parameter γ, focusing on difficult-to-classify samples.

[0055] The gradient calculation of cross-entropy is smooth, which can effectively avoid the problem of gradient disappearance in the deep UNet network. Especially in the skip connections of the encoder-decoder, the gradient can be stably backpropagated through the underlying feature path, ensuring the fusion and optimization of deep semantics and shallow details.

[0056] Adam dynamically assigns learning rates to each parameter of UNet by calculating the first and second moments of the gradient. For example, in the encoder convolutional layer, parameters with larger gradients (such as weights near the lesion edge) are given smaller learning rates to avoid oscillations, while parameters with gentle gradients (such as background regions) use larger learning rates to accelerate convergence. The skip connections of UNet require a stable gradient flow. The momentum term of Adam alleviates the problem of parameter update stagnation caused by gradient disappearance in the deep network by accumulating the historical gradient directions. To address the gradient estimation bias caused by random parameter initialization in the initial stage of UNet training, Adam corrects the first and second moments through a formula, making the learning rate adjustment more stable in the initial stage. For example, in ICG fluorescence image segmentation, this mechanism can avoid incorrect updates caused by noise interference in the first 5 epochs.

[0057] Another expression of the formula of the cross-entropy loss function is as follows:

[0058] , where, represents the true label, represents the probability value predicted by the model.

[0059] Optimizer update formula: , where, represents the model parameters, represents the learning rate, represents the loss of the current model.

[0060] Performing global segmentation on the ICG fluorescence imaging image using the trained UNet deep learning network to obtain a probability map of the watershed boundary, including:

[0061] According to the formula: , obtaining the probability map , where, represents a trained UNet deep learning network, represents the ICG fluorescence imaging image, represents the spatial coordinates, represents the time variable.

[0062] In step S130, the probability map is processed to obtain the final boundary.

[0063] The processing of the probability map to obtain the final boundary includes: performing binarization processing on the probability map to obtain a rough boundary; using the Canny edge detection algorithm to detect the ICG fluorescence imaging image to obtain a detailed boundary; fusing the rough boundary and the detailed boundary to obtain a fused boundary; and processing the fused boundary by means of a morphological operation function to obtain the final boundary.

[0064] The core objective of performing multi-level boundary processing on the ICG fluorescence imaging image is to balance the integrity of the target contour and the detail accuracy, while improving the anti-interference ability and clinical applicability of the algorithm.

[0065] The probability map identifies the target area through the confidence distribution. Binarization can discretize the continuous probability values into 0 / 1 to quickly extract the main contour. For example, in tumor detection, the high-probability area is retained as the rough boundary, filtering out low-confidence noise points to form a general morphological framework of the target. An adaptive threshold can be used to cope with the signal intensity changes caused by the injection dose and tissue depth of the ICG fluorescence.

[0066] Canny can detect low-contrast details such as microvessels in the ICG image through gradient direction interpolation and non-maximum suppression. For example, in angiography, the dual-threshold mechanism of Canny can completely extract the bifurcation structure. High-pass filtering preprocessing can eliminate shot noise and background fluorescence artifacts, and the hysteresis threshold strategy connects strong and weak edges to avoid isolated noise being misjudged as an edge.

[0067] The rough boundary provides the main framework of the target, and the detailed boundary supplements the local texture. After fusion, a boundary with both macroscopic integrity and microscopic accuracy is formed. For example, in tumor surgical navigation, the rough boundary locates the tumor theme, and the detailed boundary marks the infiltration area, reducing the risk of intraoperative residue. Fusion algorithms (such as weighted superposition or conditional random field) can dynamically adjust the weights according to the regional confidence. For example, the high-probability area is mainly based on the rough boundary, and the low-probability area relies on the Canny details to balance the robustness of the algorithm under complex illumination or tissue heterogeneity.

[0068] Morphological closing operation (first dilation then erosion) can connect the edge breaks caused by fluorescence diffusion and simultaneously eliminate the jagged artifacts in Canny detection. For example, in the optimization of blood vessel contours, the closing operation makes the edges of the vessel wall continuous and smooth. The opening operation (first erosion then dilation) can remove small noise and isolated false edges. For example, in the extraction of tumor boundaries, the opening operation filters out non-specific fluorescence signals and retains the true lesion contours. The morphological gradient operation (the difference between dilation and erosion) can enhance the edge contrast and highlight the transition region between the target and the background, facilitating subsequent quantitative analysis (such as tumor volume calculation).

[0069] As Figure 2 shown, the binarization of the probability map to obtain a rough boundary includes:

[0070] According to the formula: , the rough boundary is obtained , where represents a set threshold.

[0071] The detection of the ICG fluorescence imaging image using the Canny edge detection algorithm to obtain a detailed boundary includes:

[0072] According to the formula: , the detailed boundary is obtained .

[0073] The processing of the fusion boundary by means of a morphological operation function to obtain a final boundary includes:

[0074] According to the formula: , the final boundary is obtained , where represents the fusion boundary, represents the morphological operation function.

[0075] In step S140, sampling is performed on the final boundary to obtain a set of key point coordinates. Specifically, at least a preset number of anchor points are selected at the curvature extreme points according to the confidence threshold , and it is satisfied that the adjacent anchor points are less than or equal to , forming the set of key point coordinates , where the local curvature corresponding to the anchor point is the curvature extreme value of the corresponding pixel on the curve, the confidence of the final boundary is greater than or equal to , , and the arc length difference between adjacent anchor points along the curve is less than or equal to .

[0076] The set of key point coordinates is a set of multi-dimensional space position points composed of a group of ordered numerical values, and each coordinate point represents the position information of the target in a specific dimension, such as two-dimensional coordinates .

[0077] Adaptive sampling based on segmentation confidence, densifying sampling in the complex curve area, improves the reconstruction accuracy and reduces the redundant point bandwidth.

[0078] The UNet inference stage outputs a real-valued probability map , where the value of each pixel ranges between 0 and 1, representing the probability that the pixel belongs to the "watershed boundary". Subsequent binarization and morphological operations are only used to obtain the geometric shape, that is, a connected set of pixels, and will not destroy . The processed geometric boundary can still index the corresponding probability in .

[0079] After obtaining the final geometric boundary , for any pixel in it directly back-check . This is defined as the boundary confidence of the pixel . If the pixel is sparsely sampled due to thinning, the average value of its neighborhood probability can be taken as , to reduce the discrete error. The neighborhood average here is a fault tolerance measure for numerical robustness and is only used in scenarios where the pixel is thinned or drifted. For curvature extreme points , its is the anchor point confidence.

[0080] The confidence threshold T is used to eliminate unreliable anchor points with too low probability. The empirical value is obtained through cross-validation on 50 laparoscopic videos , and in practice, it is usually selected in the interval, higher than the binarization threshold .

[0081] In step S150, determine the displacement vector of the set of key point coordinates within a preset time interval. An improved TapNet-Siamese key point tracking network is used to determine the displacement vector of the set of key point coordinates within a preset time interval , specifically including: performing feature embedding on local image patches centered on each key point in the reference frame and the current frame through a double-branch convolutional subnetwork with shared weights; using cosine similarity to match the embedded vectors to obtain the displacement vector; when the matching confidence is lower than the threshold or the cross-frame displacement is greater than , call the cross-frame cross-attention network based on the TapNet structure to re-estimate the displacement vector, where represents the displacement threshold, denotes the confidence threshold; the TapNet-Siamese key point tracking network includes: an input layer for receiving a reference frame, a current frame, and key point coordinates, a dual-branch convolutional-TSM-ResNet18 backbone with fully shared weights, a Siamese metric learning branch for outputting key point embedding vectors, a Cross-Attention module for cross-frame feature alignment when the confidence is lower than the threshold for outputting the displacement regression head. In step S160, if the average confidence of consecutive K frames is lower than the threshold , recalculate the final boundary and refresh the set of key point coordinates .

[0082] On the basis of the original TapNet framework, an explicit Siamese (weight-sharing) metric learning branch is introduced, which improves the long-term key point re-identification and occlusion recovery capabilities. The original TapNet focuses on cross-frame feature fusion, and the backbone network structure is a dual-frame joint encoder. There are two problems in the intraoperative scenario. First, occlusion / large displacement after fluorescence attenuation: simple cross-attention is prone to mismatch. Second, the frame rate and resolution are limited: the original TapNet has the largest full-map-level interaction calculation. TapNet-Siamese supplements each key point with a local Patch descriptor. If the cross-frame displacement is greater than , or the matching confidence is less than , then Siamese Re-ID is enabled. The Siamese branch performs lightweight embedding comparison on the Patch. Only when the confidence is low does it return to the backbone for re-segmentation. Use the end-to-end network output of Siamese TapNet with a re-identification and re-localization mechanism.

[0083] TapNet-Siamese provides a tracking retention rate of greater than or equal to 95% under cross-occlusion / disparity / light drift.

[0084] The present disclosure provides a two-level fallback mechanism.

[0085] The first-level fallback mechanism: when the score of a single anchor point is less than , and the global average confidence is greater than or equal to , at the current frame, crop a pixel patch centered on this anchor point; send the patch into the lightweight sub-branch of the same UNet; take the new boundary pixels in the segmentation result for local replacement, and update the coordinates of this anchor point and . This level of fallback mechanism effectively avoids full-frame inference, takes less than 3 milliseconds, and the real-time performance is not affected. Other anchor points remain unchanged.

[0086] The second-level fallback mechanism: when consecutive The frame average confidence is lower than , indicating large-scale occlusion or overall drift. The "latest frame plus the previous frame's virtual boundary mask" is stacked into a 4-channel tensor and fed into UNet. Among them, the latest frame provides the current texture; the mask is a binary image, and the historical boundary is used as a prior to guide the network, so it can still converge. The execution result of the above actions is to produce a new probability map and the geometric boundary . For the processing of anchor points: completely discard the old anchor points, and resample according to the curvature extreme value plus confidence to obtain a new set of anchor points that satisfy , and then continue to track.

[0087] If there is a loss situation (single-point loss or multi-point continuous loss), such as the anchor point deviating from the real boundary / being occluded by the clamp resulting in a sudden drop in the Siamese similarity / obvious deformation of the surface texture, the first-level fallback mechanism (fine-grained correction) is executed for single-point loss, and the second-level fallback mechanism (steady-state recovery) is executed for multi-point continuous loss, and then resampling is performed. For example, when the continuous frame average confidence , the system stacks the current frame with the previous frame's boundary mask into a tensor, feeds it into the original UNet backbone to perform a full-frame re-segmentation; then re-generates a set of anchor points according to the curvature extreme value and criterion to ensure global consistency.

[0088] Using the latest frame can best reflect the immediate pose of the organ. If going back to the initial frame, the intraoperative deformation information will be lost, which will instead reduce the accuracy. The mask prior plus the latest frame branch can maintain .

[0089] Introduce the "non-fluorescent condition" during the training stage of the model. The training data of the model includes the near-infrared part and a small number of white-light parts. The near-infrared part is a strong fluorescence signal, and the small number of white-light parts are ordinary RGB frames. The boundary output by the near-infrared branch is used as a pseudo-label for the white-light branch, which improves the robustness and consistency expression of the model. The inputs during re-segmentation include and , is the latest frame (capturing real-time texture and deformation), is the previous frame's virtual boundary mask (shape and position prior); the two are stacked in the channel dimension to form a 4-channel tensor and fed into UNet. Even if , at this time, the white-light texture plus the prior mask can still output a probability level of about 0.3, leaving enough confidence space for subsequent anchor point screening.

[0090] In this disclosure, The closed-loop boundary of the previous moment is provided. The network guidance aims at shape similarity and texture consistency to avoid drifting to irrelevant structures. Before triggering full-frame re-segmentation, most anchor points still maintain a confidence level of 60% to 70%. The network learns that these semi-trusted points can be used as hard attention to anchor the current organ. The probability map after re-segmentation still needs to meet the anchor point rule and is smoothed by Catmull-Rom + Kalman. If the overall ..., it will be regarded as a failure and fallback to the previous stable boundary to ensure that incorrect results will not be forced on the doctor. If white light re-segmentation produces ..., the current result will be withdrawn and the historical boundary will continue to be used until a new high-confidence segmentation is generated.

[0091] In the 60 completely non-fluorescent segments (with an average of 4 minutes) of this disclosure, the boundary after re-segmentation ..., ..., compared with the scheme that only relies on tracking but does not perform re-segmentation, the delay display window is increased from 4.5 minutes to 8.0 minutes. Among them, the delay display window refers to the longest continuous time that the system can still maintain a "clinically acceptable" boundary display after the fluorescent visible signal completely decays.

[0092] Collect the complete surgical video and manually mark the fluorescence decay points ; frame by frame in subsequent frames, record the and between the virtual boundary and the offline ground truth; when the error exceeds the limit, stop timing to obtain a value for each sequence; report the median.

[0093] The method for determining the delay display window includes: record as the moment when the fluorescence signal is lower than the set brightness threshold, and at this moment, the ICG is regarded as disappeared; continuously monitor the error between the virtual boundary and the offline manually marked boundary. Specifically, use the average distance of 3 pixels or coefficient 0.75 as the standard of "still accurate". When the standard is no longer met for the first time, record the moment as . The delay display window delay = ....

[0094] In the Tap-Siamese scheme of this disclosure, delay is 8 minutes. Attached are the delays of the optical flow scheme and the original TapNet scheme, which are approximately 3 minutes and approximately 4.5 minutes respectively.

[0095] The advantages of the TapNet-Siamese key-point tracking network of the present disclosure are shown in the following table.

[0096]

[0097] In step S170, according to the set of key-point coordinates and the displacement vector, reconstruct the dynamic change of the final boundary.

[0098] The reconstructing the dynamic change of the final boundary according to the set of key points and the displacement includes:

[0099] According to the formula: , reconstruct the dynamic change of the final boundary, where represents the displacement vector of each key point within a preset time interval , and this displacement vector is obtained through an improved TapNet-Siamese key-point tracking network, represents the set of key-point coordinates sampled at time , represents the set of key-point coordinates after dynamic change .

[0100] As Figure 3 shown, it details the specific process of key-point tracking and time-lapse imaging.

[0101] First, through the set of key-point coordinates sampled from the boundary , use the key-point tracking algorithm based on the Tapnet-Siamese network to dynamically update the set of key-point coordinates. The update model can be described as:

[0102] , where represents the set of key-point coordinates sampled from the boundary at time , represents the displacement vector of each key point within the time interval , which is used to describe the movement change of the boundary in a dynamic environment. By continuously updating the key-point coordinates and accumulating the contribution of , real-time tracking of the boundary can be achieved, thereby extending the effective observation time of fluorescence imaging, that is, maintaining a stable boundary hint. Even in a complex dynamic environment where the organ flips or deforms, the continuous and stable boundary hint can still be maintained.

[0103] On the left side of the figure is a schematic diagram of the process of dynamic boundary tracking and delayed imaging, and on the right side of the figure, the actual schematic effect is shown. From left to right on the right side are the schematic pictures of each frame of the video in the time series. The first frame shows the initial moment of fluorescence imaging, in which the marked points A, B, and C are located on the fluorescence boundary. Based on the Tapnet method, the algorithm real-time tracks the features of these three points in the image and updates their positions. By the last frame, that is, the moment when the fluorescence imaging signal has disappeared, after the algorithm detects, matches, and tracks the image features of points A, B, and C, the positions of points A', B', and C' are respectively located, thus constructing a virtual fluorescence boundary. Even if the fluorescence signal disappears or the surface of the organ undergoes deformation and offset, the tracking algorithm based on image features can still achieve real-time tracking of the boundary, thereby achieving the effect of delayed imaging.

[0104] As Figure 5 shown, on the left side of the figure is the initial moment of fluorescence imaging, at this time the boundary of the target area is clearly visible. Through fluorescence segmentation and boundary extraction techniques, the target boundary is accurately identified, and the extracted virtual boundary is marked with a red line.

[0105] On the right side of the figure is the moment when the fluorescence signal has faded, and traditional fluorescence imaging can no longer provide effective boundary information. The present disclosure uses a key point tracking algorithm based on image features to real-time track the initial boundary position, and the red line is the delayed virtual boundary formed by tracking. This boundary is not statically maintained, but is dynamically detected, matched, and updated according to the image features of the organ surface, so that in complex dynamic scenarios such as respiratory movement, instrument interference, or tissue deformation, it can still accurately follow the movement of the target area with flexible boundary following.

[0106] As above, the present disclosure not only realizes the delayed display of the fluorescence boundary, but also has excellent dynamic adaptability, providing a continuous and stable visual reference for intraoperative navigation.

[0107] As Figure 4 shown, the fluorescence imaging image is collected from the endoscopic image processor, the processing unit processes the fluorescence imaging image, and finally, the display screen displays the processing result. A dedicated image processing chip (such as an FPGA) is used for boundary detection and tracking to ensure real-time processing of a large amount of fluorescence imaging images during the operation.

[0108] Figure 6 is a block diagram showing an image processing apparatus according to some embodiments of the present disclosure. As Figure 6 shown, the image processing apparatus 600 includes an image acquisition module 610, an image segmentation module 620, a probability map processing module 630, a final boundary sampling module 640, a displacement vector determination module 650, a final boundary recalculation module 660, and a final boundary reconstruction module 670.

[0109] An image acquisition module 610, configured to acquire ICG fluorescence imaging images;

[0110] An image segmentation module 620, configured to globally segment the ICG fluorescence imaging images by using a trained UNet deep learning network to obtain a probability map of the watershed boundary;

[0111] A probability map processing module 630, configured to process the probability map to obtain a final boundary;

[0112] A final boundary sampling module 640, configured to sample the final boundary to obtain a set of key point coordinates;

[0113] A displacement vector determination module 650, configured to determine a displacement vector of the set of key point coordinates within a preset time interval;

[0114] A final boundary recalculation module 660, configured to recalculate the final boundary and refresh the set of key point coordinates if the average confidence of consecutive K frames is lower than a threshold , and refresh the set of key point coordinates.

[0115] A final boundary reconstruction module 670, configured to reconstruct the dynamic changes of the final boundary according to the set of key point coordinates and the displacement vector.

[0116] In the device of the embodiments of the present disclosure, the ICG fluorescence imaging images are globally segmented by using a trained UNet deep learning network to obtain a probability map of the watershed boundary, the ICG fluorescence imaging images are detected by using a Canny edge detection algorithm to obtain a detailed boundary, and the fused boundary is processed by using a morphological operation function to obtain a final boundary, which improves the accuracy of boundary extraction in a dynamic and low signal-to-noise ratio environment. According to the set of key point coordinates and the displacement vector, the dynamic changes of the final boundary are reconstructed, which can effectively cope with the situation of organ flipping and deformation and maintain the continuous stability of the boundary prompt, that is, extend the boundary prompt time.

[0117] Figure 7 is a block diagram showing an image processing device according to some other embodiments of the present disclosure. As Figure 7 shown, the image processing device 700 includes a memory 710; and a processor 720 coupled to the memory 710. The memory 710 is used to store instructions for implementing corresponding embodiments of the image processing method. The processor 720 is configured to execute the image processing method in any of the embodiments of the present disclosure based on the instructions stored in the memory 710.

[0118] Figure 8 is a block diagram showing a computer system for implementing some embodiments of the present disclosure. As Figure 8As shown, the computer system 800 may be embodied in the form of a general-purpose computing device. The computer system 800 includes a memory 810, a processor 820, and a bus 830 that connects different system components.

[0119] The memory 810 may include, for example, a system memory, a non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs. The system memory may include a volatile storage medium, such as random access memory (RAM) and / or cache memory. The non-volatile storage medium stores, for example, instructions for corresponding embodiments that execute at least one of the image processing methods. The non-volatile storage medium includes, but is not limited to, disk memory, optical memory, flash memory, etc.

[0120] The processor 820 may be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate, or transistor and other discrete hardware components. Correspondingly, each of the modules such as an image acquisition module, an image segmentation module, a probability graph processing module, a final boundary sampling module, a displacement vector determination module, a final boundary recalculation module, and a final boundary reconstruction module may be implemented by a central processing unit (CPU) running instructions for executing corresponding steps in the memory, or may be implemented by a dedicated circuit for executing the corresponding steps.

[0121] The bus 830 may use any of a variety of bus structures. For example, the bus structure includes, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, and a Peripheral Component Interconnect (PCI) bus.

[0122] The computer system 800 may further include an input / output interface 840, a network interface 850, a storage interface 860, etc. These interfaces 840, 850, 860, as well as the memory 810 and the processor 820, may be connected via the bus 830. The input / output interface 840 provides a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 850 provides a connection interface for various networking devices. The storage interface 860 provides a connection interface for external storage devices such as floppy disks, USB flash drives, and SD cards.

[0123] Here, various aspects of the present disclosure have been described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of the blocks, may be implemented by computer-readable program instructions.

[0124] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable apparatus to produce a machine, such that the instructions executed by the processor create means for implementing the functions specified in one or more boxes in the flowchart and / or block diagram.

[0125] These computer-readable program instructions may also be stored in a computer-readable memory, which instructions cause a computer to operate in a particular manner, thereby producing a manufacture including instructions for implementing the functions specified in one or more boxes in the flowchart and / or block diagram.

[0126] The present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0127] In the present disclosure, a trained UNet deep learning network is used to globally segment ICG fluorescence imaging images to obtain a probability map of the watershed boundary. The Canny edge detection algorithm is used to detect the ICG fluorescence imaging images to obtain the detail boundary. The fused boundary is processed by means of a morphological operation function to obtain the final boundary, improving the accuracy of boundary extraction in a dynamic and low signal-to-noise ratio environment. According to the set of key point coordinates and the displacement vector, the dynamic changes of the final boundary are reconstructed, which can effectively cope with the flipping and deformation of organs and maintain the continuous stability of the boundary hint, that is, extend the boundary hint time.

[0128] Thus far, the image processing method and apparatus according to the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details well known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

[0129] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining an ICG fluorescence imaging image; Using a trained UNet deep learning network to globally segment the ICG fluorescence imaging image to obtain a probability map of the watershed boundary; Processing the probability map to obtain a final boundary; Sample the final boundary to obtain a set of key point coordinates, specifically including: at the curvature extreme points, according to the confidence threshold Select no less than the preset number of anchor points, and satisfy that the adjacent anchor points are less than or equal to , to form the set of key point coordinates. Among them, the local curvature corresponding to the anchor point is the curvature extreme value of the corresponding pixel on the curve, and the confidence of the final boundary is greater than or equal to , , the arc length difference between adjacent anchor points along the curve is less than or equal to ; Use the improved TapNet-Siamese key-point tracking network to determine the displacement vectors of the set of key-point coordinates within a preset time interval Specifically, it includes: performing feature embedding on local image patches centered on each key point in the reference frame and the current frame through a double-branch convolutional sub-network with shared weights; obtaining the displacement vectors by matching the embedded vectors using cosine similarity; when the matching confidence is lower than the threshold or the cross-frame displacement is greater than , calling the cross-frame cross-attention network based on the TapNet structure to re-estimate the displacement vectors, where represents the displacement threshold; the TapNet-Siamese key-point tracking network includes: an input layer for receiving the reference frame, the current frame, and the key-point coordinates, a double-branch convolutional-TSM-ResNet18 backbone with fully shared weights, a Siamese metric learning branch for outputting key-point embedding vectors, a Cross-Attention module for cross-frame feature alignment when the confidence is lower than the threshold , and a displacement regression head for outputting displacement vectors; If the average confidence of consecutive K frames is lower than the threshold , recalculate the final boundary and refresh the set of key point coordinates; Reconstructing the dynamic change of the final boundary according to the set of key point coordinates and the displacement vector.

2. The image processing method according to claim 1, wherein The processing of the probability map includes: Performing binarization processing on the probability map to obtain a rough boundary; Using the Canny edge detection algorithm to detect the ICG fluorescence imaging image to obtain a detailed boundary; Fusing the rough boundary and the detailed boundary to obtain a fused boundary; Processing the fused boundary by means of a morphological operation function to obtain a final boundary.

3. The image processing method according to claim 1, characterized in that, Before the obtaining of the ICG fluorescence imaging image, it further includes: Collecting an initial ICG fluorescence imaging image; Performing grayscale processing on the initial ICG fluorescence imaging image to obtain a grayscale image; Removing noise from the grayscale image to obtain an ICG fluorescence imaging image; Returning the ICG fluorescence imaging image.

4. The image processing method according to claim 1, wherein When using the trained UNet deep learning network to globally segment the ICG fluorescence imaging image to obtain a probability map of the watershed boundary, the cross-entropy loss function is used to optimize the UNet deep learning network, and the Adam optimizer is used to update the parameters of the UNet deep learning network.

5. The image processing method according to claim 2, wherein The using of the trained UNet deep learning network to globally segment the ICG fluorescence imaging image to obtain a probability map of the watershed boundary includes: According to the formula: , the probability map is obtained, where represents the trained UNet deep learning network, represents the ICG fluorescence imaging image, represents the spatial coordinates, represents the time variable.

6. The image processing method according to claim 5, wherein The performing of the binarization processing on the probability map to obtain a rough boundary includes: According to the formula: , the rough boundary is obtained, where represents the set threshold value.

7. The image processing method according to claim 6, characterized in that The using of the Canny edge detection algorithm to detect the ICG fluorescence imaging image to obtain a detailed boundary includes: According to the formula: , the detailed boundary is obtained .

8. The image processing method according to claim 7, wherein The processing of the fused boundary by means of a morphological operation function to obtain a final boundary includes: According to the formula: , the final boundary is obtained, where represents the fusion boundary, represents the morphological operation function.

9. The image processing method according to claim 8, wherein The reconstructing of the dynamic change of the final boundary according to the set of key point coordinates and the displacement vector includes: According to the formula: , reconstruct the dynamic change of the final boundary, where represents the displacement vectors of each key point within the preset time interval , represents the set of key point coordinates sampled at time , represents the set of key point coordinates after the dynamic change .

10. An image processing apparatus, characterized in that, Includes: A memory; And A processor coupled to the memory, the processor being configured to execute the image processing method according to any one of claims 1 to 9 based on instructions stored in the memory.

Citation Information

Patent Citations

  • Fluorescence navigation system used in tumor surgery

    CN101770141A

  • Fluorescence imaging device applying spectrum detection and testing method

    CN112414987A

  • Structured landmark detection via topology-adapting deep graph learning

    US20230245329A1

  • Methods and systems for identifying lymph nodes using fluorescence imaging data

    US20250031971A1