Image processing method and device

The ICG fluorescence imaging images are processed through deep learning networks and edge detection algorithms, extracting and reconstructing the boundaries in laparoscopic surgery, solving the problem of short observation time window in ICG fluorescence imaging technology, and improving the accuracy of boundary calibration and dynamic adaptability.

CN120107301AActive Publication Date: 2025-06-06CHONGQING FUDIMAI DIGITAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510584722.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The short observation time window in ICG fluorescence imaging technology leads to limited accuracy of boundary calibration in laparoscopic surgery, especially in complex surgical environments.

Method used

The trained UNet deep learning network is used to perform global segmentation of ICG fluorescence imaging images, combined with Canny edge detection algorithm and morphological operations, the final boundary is extracted and reconstructed, and the dynamic changes of boundaries are determined through the TapNet-Siamese key point tracking network.

Benefits of technology

It improves the accuracy of boundary extraction in dynamic, low signal-to-noise ratio environments, extends the boundary prompt time, and ensures the accuracy of boundary calibration in complex laparoscopic surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107301A_ABST
    Figure CN120107301A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, and belongs to the technical field of data processing. The method comprises the following steps: acquiring an ICG fluorescence imaging image; performing global segmentation on the ICG fluorescence imaging image by using a trained UNet deep learning network to obtain a probability graph of a drainage basin boundary; processing the probability graph to obtain a final boundary; sampling the final boundary to obtain a key point coordinate set; determining a displacement vector of the key point coordinate set in a preset time interval; and reconstructing the dynamic change of the final boundary according to the key point coordinate set and the displacement vector. According to the invention, the delayed display of the fluorescence boundary is realized, and the dynamic adaptability is excellent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image processing method and device. Background Art

[0002] With the development of minimally invasive surgical techniques, laparoscopic surgery is becoming increasingly popular in clinical applications as a precise and low-invasive treatment method. During laparoscopic surgery, how to accurately demarcate the boundaries of organs, tumors or other anatomical structures is crucial to improving the safety of surgery and the therapeutic effect.

[0003] Fluorescence imaging technology, especially the use of ICG (indocyanine green) fluorescent dye for boundary calibration, has become an important means of intraoperative boundary identification.

[0004] ICG fluorescence imaging technology has good imaging effects, but its optimal observation time window is short. The development time of ICG is usually 35 seconds, approximately between 29 seconds and 44 seconds, and its development duration is about 3 minutes. The diffusion of ICG in the microcirculation will cause the watershed boundaries to gradually blur. Therefore, the optimal observation window is usually within 30 seconds to 1 minute after staining. The short time window poses certain challenges to precise operations during surgery, especially in complex laparoscopic surgery, which requires marking of watershed boundaries in a very short time to ensure the accuracy of the surgery.

[0005] The Chinese patent document with publication number CN112414987A discloses a technology named "a fluorescence imaging device and test method using spectral detection". This technology cannot accurately extract boundary details in the fluorescence imaging window, and it is difficult to solve the problem of blurred or disappeared boundaries. Existing boundary detection technologies mostly rely on traditional image processing methods, which lack sufficient robustness and accuracy when facing dynamic laparoscopic fields and complex tissue structures. The Chinese patent document with publication number CN101770141A discloses a technology named "a fluorescence navigation system for tumor surgery". Although this technology can achieve boundary tracking to a certain extent, it has poor adaptability to dynamic changes such as organ flipping and deformation, and cannot ensure the continuous stability of boundary prompts. Summary of the invention

[0006] The present disclosure proposes an image processing method and device, aiming to solve the problem of short observation time window in ICG fluorescence imaging technology.

[0007] According to a first aspect of the present disclosure, an image processing method is provided, the method comprising: obtaining an ICG fluorescence imaging image; performing global segmentation on the ICG fluorescence imaging image using a trained UNet deep learning network to obtain a probability map of a watershed boundary; processing the probability map to obtain a final boundary; sampling the final boundary to obtain a set of key point coordinates , specifically including: at the extreme value of curvature according to the confidence threshold Select no less than the preset number of anchor points, and the adjacent anchor points are less than or equal to , forming the key point coordinate set , where the local curvature corresponding to the anchor point is the curvature extreme value of the corresponding pixel on the curve, and the final boundary confidence is greater than or equal to , , the difference in arc lengths of adjacent anchor points along the curve is less than or equal to ; Use the improved TapNet-Siamese key point tracking network to determine the key point coordinate set At preset time intervals The displacement vector inside ,Right now , specifically including: embedding features of local image blocks centered on each key point in the reference frame and the current frame through a weight-sharing two-branch convolutional subnetwork; matching the embedding vectors using cosine similarity to obtain the displacement vector; when the matching confidence is lower than a threshold Or the cross-frame displacement is greater than , the cross-frame cross-attention network based on the TapNet structure is called to re-estimate the displacement vector, where: represents the displacement threshold, Represents the confidence threshold; the TapNet-Siamese key point tracking network includes: an input layer for receiving the reference frame, the current frame, and the key point coordinates, a dual convolution-TSM-ResNet18 backbone with fully shared weights, a Siamese metric learning branch for outputting the key point embedding vector, and a Cross-Attention module for cross-frame feature alignment, used to output If the average confidence of consecutive K frames is lower than the threshold , recalculate the probability map and refresh the key point coordinate set ; According to the key point coordinate set Together with the displacement vector, the dynamic changes of the final boundary are reconstructed.

[0008] In some embodiments, the processing of the probability map to obtain the final boundary includes: binarizing the probability map to obtain a rough boundary; detecting the ICG fluorescence imaging image using a Canny edge detection algorithm to obtain a detailed boundary; fusing the rough boundary with the detailed boundary to obtain a fused boundary; and processing the fused boundary with the aid of a morphological operation function to obtain a final boundary.

[0009] In some embodiments, before acquiring the ICG fluorescence imaging image, it also includes: acquiring an initial ICG fluorescence imaging image; graying the initial ICG fluorescence imaging image to obtain a grayscale image; removing noise from the grayscale image to obtain an ICG fluorescence imaging image; and returning the ICG fluorescence imaging image.

[0010] In some embodiments, the ICG fluorescence imaging image is globally segmented using a trained UNet deep learning network to obtain a probability map of the watershed boundary, wherein the UNet deep learning network is optimized using a cross entropy loss function, and the parameters of the UNet deep learning network are updated using an Adam optimizer.

[0011] In some embodiments, the use of a trained UNet deep learning network to globally segment the ICG fluorescence imaging image to obtain a probability map of the watershed boundary includes: according to the formula: , and get the probability map ,in, represents the trained UNet deep learning network, represents the ICG fluorescence imaging image, represents the spatial coordinates, Represents a time variable.

[0012] In some embodiments, binarizing the probability map to obtain a rough boundary includes: according to the formula: , and get a rough boundary ,in, Indicates setting a threshold.

[0013] In some embodiments, the use of the Canny edge detection algorithm to detect the ICG fluorescence imaging image to obtain detail boundaries includes: according to the formula: , get the detail boundary .

[0014] In some embodiments, the processing of the fused boundary by means of a morphological operation function to obtain a final boundary includes: according to the formula: , and obtain the final boundary ,in, represents the fusion boundary, Represents morphological operation function.

[0015] In some embodiments, the reconstructing the dynamic change of the final boundary according to the key point coordinate set and the displacement vector includes: according to the formula: , reconstructing the dynamic change of the final boundary, where Indicates the preset time interval The displacement vectors of each key point in Indicates at time The sampled key point coordinate set, Represents the key point coordinate set after dynamic changes .

[0016] According to a second aspect of the present disclosure, there is provided an image processing apparatus, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the above-mentioned image processing method based on instructions stored in the memory.

[0017] By adopting the above technical solution, the embodiments of the present disclosure can achieve the following beneficial technical effects: using the trained UNet deep learning network to perform global segmentation on the ICG fluorescence imaging image, obtaining the probability map of the watershed boundary, using the Canny edge detection algorithm to detect the ICG fluorescence imaging image, obtaining the detail boundary, and using the morphological operation function to process the fused boundary to obtain the final boundary, thereby improving the accuracy of boundary extraction in dynamic and low signal-to-noise ratio environments. According to the key point coordinate set and displacement vector, the dynamic changes of the final boundary are reconstructed, which can effectively deal with the situation of organ flipping and deformation, and maintain the continuous stability of the boundary prompt, that is, extend the boundary prompt time.

[0018] This paper adds a new key point screening rule: select N anchor points according to the local maximum curvature of the final boundary plus the confidence threshold, and N is greater than or equal to 50, and the difference in arc length between adjacent anchor points is less than or equal to ,Traditional optical flow does not involve adaptive sampling of segmentation ,confidence, and encrypts sampling in complex areas of curves, which improves ,reconstruction accuracy and reduces the bandwidth of redundant points.

[0019] This paper introduces an explicit Siamese (weight sharing) metric learning branch based on the original TapNet framework, which improves the long-term key point re-identification and occlusion recovery capabilities; the original TapNet focuses on cross-frame feature fusion, and the backbone network structure is a dual-frame joint encoder (TSM-ResNet18+Cross-Attention). There are two problems in the intraoperative scene, namely, occlusion / large displacement after fluorescence attenuation: simple cross-attention is prone to mismatch, and frame rate plus resolution are limited: the original TapNet has a large amount of full-image interactive calculations. This paper proposes a TapNet-Siamese key point tracking network, which adds a local Patch descriptor to each key point; if the cross-frame is greater than or the match confidence is less than , then Siamese Re-ID is enabled; the Siamese branch performs lightweight embedding comparison on the patch; only when the confidence is low does it return to the trunk for heavy segmentation. It is specified that the end-to-end network using Siamese TapNet is output And with re-identification and re-localization mechanism (score is less than Compared with optical flow (local gradient method) and TapNet (global metric learning), the present invention provides a tracking retention rate of greater than or equal to 95% across occlusion / parallax / illumination drift.

[0020] In the present disclosure, if the average anchor point confidence for consecutive K frames is less than , it automatically triggers re-segmentation, introduces a closed-loop self-recovery mechanism, and the delayed display window is extended from 3 minutes to more than 8 minutes. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0022] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings.

[0023] Figure 1 is a flowchart illustrating an image processing method according to some embodiments of the present disclosure.

[0024] Figure 2 FIG. 4 is a flow chart showing a boundary extraction process according to some embodiments of the present disclosure.

[0025] Figure 3 is a schematic diagram showing dynamically changing boundary reconstruction according to some embodiments of the present disclosure.

[0026] Figure 4 is a schematic diagram showing an image display system according to some embodiments of the present disclosure.

[0027] Figure 5 It is a schematic diagram showing the time-lapse display effect of a fluorescent imaging image according to some embodiments of the present disclosure.

[0028] Figure 6 is a block diagram illustrating an image processing apparatus according to some embodiments of the present disclosure.

[0029] Figure 7 is a block diagram showing an image processing apparatus according to some other embodiments of the present disclosure.

[0030] Figure 8 is a block diagram illustrating a computer system for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0031] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.

[0032] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0033] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0034] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.

[0035] In all examples shown and discussed herein, any specific values ​​should be understood as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0036] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0037] At present, with the development of minimally invasive surgical techniques, laparoscopic surgery is becoming increasingly popular in clinical applications as a precise and low-invasive treatment method. During laparoscopic surgery, how to accurately demarcate the boundaries of organs, tumors or other anatomical structures is crucial to improving the safety of surgery and the therapeutic effect.

[0038] Fluorescence imaging technology, especially the use of ICG (indocyanine green) fluorescent dye for boundary calibration, has become an important means of intraoperative boundary identification.

[0039] ICG fluorescence imaging technology has good imaging effects, but its optimal observation time window is short. The development time of ICG is usually 35 seconds, approximately between 29 seconds and 44 seconds, and its development duration is about 3 minutes. The diffusion of ICG in the microcirculation will cause the watershed boundaries to gradually blur. Therefore, the optimal observation window is usually within 30 seconds to 1 minute after staining. The short time window poses certain challenges to precise operations during surgery, especially in complex laparoscopic surgery, which requires marking of watershed boundaries in a very short time to ensure the accuracy of the surgery.

[0040] In view of this, the present disclosure proposes an image processing method and device, which uses a trained UNet deep learning network to perform global segmentation on the ICG fluorescence imaging image to obtain a probability map of the watershed boundary, uses the Canny edge detection algorithm to detect the ICG fluorescence imaging image to obtain the detail boundary, and uses the morphological operation function to process the fused boundary to obtain the final boundary, thereby improving the accuracy of boundary extraction in dynamic and low signal-to-noise ratio environments. According to the key point coordinate set and displacement vector, the dynamic changes of the final boundary are reconstructed, which can effectively deal with the situation of organ flipping and deformation, and maintain the continuous stability of the boundary prompt, that is, extend the boundary prompt time.

[0041] Figure 1 FIG. 4 is a flowchart showing an image processing method according to some embodiments of the present disclosure. Figure 1 As shown, the image processing method includes steps S110 to S170.

[0042] In step S110, an ICG fluorescence imaging image is acquired.

[0043] ICG fluorescence imaging technology is a near-infrared optical imaging method based on indocyanine green dye. It provides precise navigation for surgery by displaying the fluorescence signals of tissues or lesions in real time.

[0044] Before acquiring the ICG fluorescence imaging image, the method also includes: acquiring an initial ICG fluorescence imaging image; graying the initial ICG fluorescence imaging image to obtain a grayscale image; removing noise from the grayscale image to obtain an ICG fluorescence imaging image; and returning the ICG fluorescence imaging image.

[0045] like Figure 2 As shown, according to the formula: , grayscale processing is performed, where Indicates grayscale processing.

[0046] According to the formula: , to remove noise, where Indicates noise removal.

[0047] Image grayscale processing is the process of converting a color image into a grayscale image, that is, converting the red, green, and blue color channel values ​​of each pixel into a single grayscale value through mathematical methods. The grayscale value range of each pixel in a grayscale image is 0-255, where 0 is pure black and 255 is pure white. By eliminating color information and retaining only brightness information, the core formula of grayscale can be expressed as: ,in, Represents the conversion function. Different methods correspond to different calculation logics. The purpose of grayscale processing is to simplify data, enhance contrast, and improve efficiency. That is, the grayscale image only needs to be stored in a single channel, and the amount of data is reduced to 1 / 3 of the original image. Remove color interference and highlight the image structure to enhance contrast. Reduce the computational complexity of subsequent processing, such as edge detection and image segmentation.

[0048] The main purpose of denoising is to restore the real information of the image by eliminating or suppressing noise interference, and to improve the reliability of subsequent image processing and analysis. For example, eliminating high-frequency noise interference allows the Canny algorithm to locate edges more accurately. For example, suppressing artifacts caused by noise improves the robustness of the segmentation algorithm.

[0049] In step S120, the ICG fluorescence imaging image is globally segmented using the trained UNet deep learning network to obtain a probability map of the watershed boundary.

[0050] The UNet deep learning network performs global segmentation of ICG fluorescence imaging images, aiming to solve key clinical needs in medical image segmentation through high-precision, real-time pixel-level classification. Taking tumor boundary recognition as an example, ICG will retain fluorescence in the tumor area due to metabolic abnormalities. UNet extracts multi-scale features through the encoder-decoder structure, combines jump connections to fuse shallow details with deep semantic information, and can accurately segment tumor boundaries, especially in low signal-to-noise ratio scenarios.

[0051] The ICG fluorescence imaging image is globally segmented using the trained UNet deep learning network to obtain a probability map of the watershed boundary, wherein the UNet deep learning network is optimized using a cross entropy loss function, and the parameters of the UNet deep learning network are updated using an Adam optimizer.

[0052] The main functions of the cross entropy function in UNet include pixel-level classification optimization, category imbalance processing, gradient characteristics and model convergence.

[0053] The core goal of UNet is to perform pixel-level segmentation on images. Cross entropy drives the update of network parameters by quantifying the difference between the predicted probability distribution and the true label distribution of each pixel. Its mathematical expression is: ,in, represents the true label, Represents the class probability predicted by the model. This loss function is particularly suitable for binary classification (such as lesion / normal tissue) or multi-classification scenarios.

[0054] The proportion of lesion areas in medical images is usually small. Cross entropy alleviates class imbalance through the following strategies: First, weighted cross entropy: give higher weights to minority classes, such as increasing the loss contribution of lesion pixels in tumor segmentation. Second, focus loss: reduce the weight of easy-to-classify samples by adjusting the parameter Y, and focus on difficult-to-classify samples.

[0055] The gradient calculation of cross entropy is smooth, which can effectively avoid the gradient vanishing problem in the UNet deep network. Especially in the jump connection of the encoder-decoder, the gradient can be stably transmitted back through the underlying feature path to ensure the fusion optimization of deep semantics and shallow details.

[0056] Adam dynamically assigns a learning rate to each parameter of UNet by calculating the first-order and second-order moments of the gradient. For example, in the encoder convolution layer, parameters with larger gradients (such as weights close to the edge of the lesion) are assigned smaller learning rates to avoid oscillations, while parameters with gentle gradients (such as background areas) use larger learning rates to accelerate convergence. UNet's jump connections require a stable gradient flow, and Adam's momentum term alleviates the problem of parameter update stagnation caused by gradient vanishing in deep networks by accumulating historical gradient directions. In response to the gradient estimation bias caused by random initialization of parameters in the early stage of UNet training, Adam corrects the first-order and second-order moments through formulas to make the initial stage of learning rate adjustment more stable. For example, in ICG fluorescence image segmentation, this mechanism can avoid erroneous updates caused by noise interference in the first 5 epochs.

[0057] Another expression of the formula for the cross entropy loss function is as follows: ,in, represents the true label, Represents the probability value predicted by the model.

[0058] Optimizer update formula: ,in, represents the model parameters, represents the learning rate, Represents the loss of the current model.

[0059] The ICG fluorescence imaging image is globally segmented using the trained UNet deep learning network to obtain a probability map of the watershed boundary, including: According to the formula: , and get the probability map ,in, represents the trained UNet deep learning network, represents the ICG fluorescence imaging image, represents the spatial coordinates, Represents a time variable.

[0060] In step S130, the probability map is processed to obtain a final boundary.

[0061] The processing of the probability map to obtain the final boundary includes: binarizing the probability map to obtain a rough boundary; detecting the ICG fluorescence imaging image using a Canny edge detection algorithm to obtain a detailed boundary; fusing the rough boundary with the detailed boundary to obtain a fused boundary; and processing the fused boundary with the aid of a morphological operation function to obtain a final boundary.

[0062] The core goal of multi-level boundary processing of ICG fluorescence imaging images is to balance the integrity and detail accuracy of the target contour, while improving the algorithm's anti-interference ability and clinical applicability.

[0063] The probability map identifies the target area through the confidence distribution, and the binarization can discretize the continuous probability value into 0 / 1 to quickly extract the main contour. For example, in tumor detection, the high probability area is retained as a rough boundary, and the low confidence noise points are filtered to form a rough morphological framework of the target. The use of adaptive thresholds can cope with the signal intensity changes of ICG fluorescence caused by injection dose and tissue depth.

[0064] Canny can detect low-contrast details such as microvessels in ICG images through gradient interpolation and non-maximum suppression. For example, in angiography, Canny's dual threshold mechanism can fully extract bifurcation structures. High-filter preprocessing can eliminate shot noise and background fluorescence artifacts, while the hysteresis threshold strategy connects strong and weak edges to avoid isolated noise being misjudged as edges.

[0065] The rough boundary provides the main framework of the target, and the detailed boundary supplements the local texture. After fusion, a boundary with both macroscopic integrity and microscopic precision is formed. For example, in tumor surgical navigation, the rough boundary locates the tumor theme, and the detailed boundary marks the infiltration area to reduce the risk of residual during surgery. Fusion algorithms (such as weighted superposition or conditional random fields) can dynamically adjust weights according to regional confidence. For example, high-probability areas are mainly based on rough boundaries, and low-probability areas rely on Canny details, balancing the robustness of the algorithm under complex lighting or tissue heterogeneity.

[0066] The morphological closing operation (dilation followed by erosion) can connect edge breaks caused by fluorescence diffusion and eliminate jagged artifacts in Canny detection. For example, in vascular contour optimization, the closing operation makes the edge of the tube wall continuous and smooth. The opening operation (erosion followed by dilation) can remove tiny noise points and isolated false edges. For example, in tumor boundary extraction, the opening operation filters out non-specific fluorescence signals and retains the true lesion contour. The morphological gradient operation (difference between dilation and erosion) can enhance edge contrast and highlight the transition area between the target and the background, which is convenient for subsequent quantitative analysis (such as tumor volume calculation).

[0067] like Figure 2 As shown, the probability map is binarized to obtain a rough boundary, including: According to the formula: , and get a rough boundary ,in, Indicates setting a threshold.

[0068] The method of detecting the ICG fluorescence imaging image using the Canny edge detection algorithm to obtain detail boundaries includes: According to the formula: , get the detail boundary .

[0069] The step of processing the fused boundary by means of a morphological operation function to obtain a final boundary includes: According to the formula: , and obtain the final boundary ,in, represents the fusion boundary, Represents morphological operation function.

[0070] In step S140, the final boundary is sampled to obtain a set of key point coordinates. Specifically, the following steps are performed: sampling the final boundary at the curvature extreme value according to the confidence threshold. Select no less than the preset number of anchor points, and the number of adjacent anchor points is less than or equal to , forming the key point coordinate set , where the local curvature corresponding to the anchor point is the curvature extreme value of the corresponding pixel on the curve, and the final boundary confidence is greater than or equal to , , the difference in arc lengths of adjacent anchor points along the curve is less than or equal to .

[0071] Keypoint coordinate set It is a set of multidimensional space location points composed of a set of ordered values. Each coordinate point represents the location information of the target in a specific dimension, such as the two-dimensional coordinate .

[0072] Adaptive sampling based on segmentation confidence densifies sampling in complex areas of the curve, improves reconstruction accuracy, and reduces redundant point bandwidth.

[0073] The UNet inference phase outputs a real-valued probability map , the value of each pixel is between 0 and 1, indicating the probability that the pixel belongs to the "watershed boundary". The subsequent binarization and morphological operations are only used to obtain the geometric shape, that is, the connected pixel set, and will not destroy The processed geometric boundaries can still be found in The indices correspond to probabilities.

[0074] After obtaining the final geometric boundary , for any pixel Directly check back This is defined as the pixel boundary confidence If the pixel is sparsely sampled due to thinning, you can take its The average value of the domain probability is , in order to reduce the discrete error. Here The neighborhood average is a fault-tolerant method for numerical robustness and is only used when pixels are thinned or drifting. ,That This is the anchor point confidence.

[0075] The confidence threshold T is used to remove unreliable anchor points with too low probability, and the empirical value is obtained by cross-validation on 50 laparoscopic videos. In practice, it is usually chosen Interval, above the binarization threshold .

[0076] In step S150, determine the key point coordinate set The displacement vector within a preset time interval. The improved TapNet-Siamese key point tracking network is used to determine the key point coordinate set At preset time intervals The displacement vector inside , specifically including: embedding features of local image blocks centered on each key point in the reference frame and the current frame through a weight-sharing two-branch convolutional subnetwork; matching the embedding vectors using cosine similarity to obtain the displacement vector; when the matching confidence is lower than a threshold Or the cross-frame displacement is greater than , the cross-frame cross-attention network based on the TapNet structure is called to re-estimate the displacement vector, where: represents the displacement threshold, Represents the confidence threshold; the TapNet-Siamese key point tracking network includes: an input layer for receiving the reference frame, the current frame, and the key point coordinates, a dual convolution-TSM-ResNet18 backbone with fully shared weights, a Siamese metric learning branch for outputting the key point embedding vector, and a Cross-Attention module for cross-frame feature alignment, used to output In step S160, if the average confidence of consecutive K frames is lower than the threshold , recalculate the final boundary and refresh the key point coordinate set .

[0077] Based on the original TapNet framework, an explicit Siamese (weight sharing) metric learning branch is introduced to improve the long-term key point re-identification and occlusion recovery capabilities. The original TapNet focuses on cross-frame feature fusion, and the backbone network structure is a dual-frame joint encoder. There are two problems with intraoperative scenes. First, occlusion / large displacement after fluorescence attenuation: simple cross-attention is prone to mismatch. Second, frame rate and resolution are limited: the original TapNet has the largest full-image interactive calculation. TapNet-Siamese adds a local Patch descriptor to each key point. If the cross-frame displacement is greater than , or the match confidence is less than , Siamese Re-ID is enabled. The Siamese branch performs lightweight embedding comparison on the patch. It only returns to the main trunk for re-segmentation when the confidence is low. The end-to-end network output using Siamese TapNet And with re-identification and relocation mechanism.

[0078] TapNet-Siamese provides a tracking retention rate of greater than or equal to 95% across occlusion / parallax / illumination drift.

[0079] The present disclosure provides a two-level fallback mechanism.

[0080] First-level fallback mechanism: When a single anchor point score is less than , and the global average confidence is greater than or equal to When , crop with the anchor point as the center in the current frame Pixel patch; send the patch to the light quantum branch of the same UNet; take the new boundary pixels in the segmentation result for local replacement, update the coordinates of the anchor point and ,This level of fallback mechanism effectively avoids full frame inference, takes less than 3 milliseconds, the real time is not affected, and other anchor points remain unchanged.

[0081] Second level fallback mechanism: When continuous The frame average confidence is lower than , indicating large-scale occlusion or overall drift, the "latest frame plus a frame of virtual boundary mask" is stacked into a 4-channel tensor and sent to UNet, where the latest frame provides the current texture; the mask is a binary image, and the historical boundary is used as a priori to guide the network, so it can still converge. The execution result of the above actions is to produce a new probability map With geometric boundaries The processing of anchor points is as follows: completely discard the old anchor points, add confidence according to the extreme value of curvature Resampling satisfies The new anchor point set is used to continue tracking.

[0082] If there is a loss situation (single point loss or multiple points continuous loss), such as the anchor point is out of the real boundary / occluded by the clamp, resulting in a sudden drop in Siamese similarity / obvious deformation of the surface texture, the single point loss executes the first level fallback mechanism (fine-grained correction), and the multiple points continuous loss executes the second level fallback mechanism (steady-state recovery), and then resamples. For example, when the continuous Frame average confidence When the system sets the current frame Boundary mask with previous frame Stacked as Tensor, sent to the original UNet backbone to perform a full-frame re-segmentation; then based on the curvature extreme value and The anchor point set is regenerated based on the criterion to ensure global consistency.

[0083] Using the latest frame can best reflect the immediate position of the organ. If we return to the initial frame, we will lose the intraoperative deformation information and reduce the accuracy. Mask prior plus the latest frame branch can maintain .

[0084] The "no fluorescence condition" is introduced during the training phase of the model. The training data of the model includes the near-infrared part and a small amount of white light. The near-infrared part is a strong fluorescence signal, and the small amount of white light is a normal RGB frame. The boundary produced by the near-infrared branch is used as a pseudo-label for the white light branch, which improves the robustness and consistency of the model. The input during re-segmentation includes and , for the latest frame (capturing real-time texture and deformation), is the virtual boundary mask of the previous frame (shape and position prior); the two are stacked in the channel dimension to form a 4-channel tensor and fed into UNet. At this time, the white light texture plus the prior mask can still output a probability level of about 0.3, leaving enough confidence space for subsequent anchor point screening.

[0085] In this disclosure, The closed-loop boundary of the previous moment is provided, and the network is guided to aim at shape similarity and texture consistency to avoid drifting to irrelevant structures; before triggering full-frame re-segmentation, most anchor points still maintain a confidence level of 60% to 70%; the network learns that these semi-credible points can be used as hard attention anchors for the current organ; the probability map after re-segmentation still needs to meet Anchor point rule, and smoothed by Catmull-Rom+Kalman; if the overall , will be considered a failure and fall back to the previous stable boundary to ensure that the wrong result will not be forced on the doctor. , then the current result is withdrawn and the historical boundary continues to be used until a new high-confidence segmentation is generated.

[0086] The present invention discloses the boundary after re-segmentation on 60 segments without fluorescence (average 4 minutes). , Compared with the solution that only relies on tracking but not re-segmentation, the delayed display window is increased from 4.5min to 8.0min, where the delayed display window refers to the longest continuous time that the system can maintain a "clinically acceptable" boundary display after the fluorescent visible signal is completely attenuated.

[0087] Collect the complete surgical video and manually mark the fluorescence attenuation points ; In subsequent frames, record the virtual boundary and offline true value frame by frame and ; When the error exceeds the limit, stop timing and get Extend each series by one value; report the median.

[0088] The method for determining the delayed display window includes: The moment when the fluorescence signal is lower than the set brightness threshold, ICG is considered to have disappeared at this moment; the error between the virtual boundary and the offline manually marked boundary is continuously monitored. Specifically, the average distance 3 pixels or coefficient 0.75 is used as the "still accurate" standard. The moment when this standard is no longer met for the first time is recorded as . Delayed display window Extension = .

[0089] In the Tap-Siamese scheme of the present disclosure, Extension 8min. Attached is the optical flow solution and the original TapNet solution The delay is about 3 minutes and about 4.5 minutes respectively.

[0090] The advantages of the TapNet-Siamese key point tracking network disclosed in the present invention are shown in the following table.

[0091]

[0092] In step S170, according to the key point coordinate set and the displacement vector, to reconstruct the dynamic change of the final boundary.

[0093] According to the key point set and the displacement, reconstructing the dynamic change of the final boundary, including: According to the formula: , reconstructing the dynamic change of the final boundary, where Indicates the preset time interval The displacement vector of each key point in is obtained by the improved TapNet-Siamese key point tracking network. Indicates at time The sampled key point coordinate set, Represents the key point coordinate set after dynamic changes .

[0094] like Figure 3 As shown, the specific process of key point tracking and time-lapse imaging is demonstrated in detail.

[0095] First, by The key point coordinate set sampled from , the key point tracking algorithm based on Tapnet-Siamese network is used to dynamically update the key point coordinate set. The update model can be described as: ,in, Indicates at time From the border The set of key point coordinates obtained by sampling, Indicates the time interval The displacement vectors of each key point in the boundary are used to describe the movement of the boundary in a dynamic environment. By continuously updating the coordinates of the key points , and accumulate The contribution of the sensor can achieve real-time tracking of the boundary, thereby extending the effective observation time of fluorescence imaging, that is, maintaining a stable boundary prompt. Even in a complex dynamic environment where the organ flips or deforms, it can still maintain a continuous and stable boundary prompt.

[0096] The left part of the figure is a schematic diagram of the process of dynamic boundary tracking and time-lapse imaging, and the right part of the figure shows the actual schematic effect. The right side shows schematic diagrams of each frame of the video in the time series from left to right. The first frame shows the initial moment of fluorescence imaging, in which the marked points A, B, and C are located on the fluorescence boundary. Based on the Tapnet method, the algorithm tracks the features of these three points in the image in real time and updates their positions. In the last frame, when the fluorescence imaging signal has disappeared, after the algorithm detects, matches, and tracks the image features of points A, B, and C, it locates the positions of points A′, B′, and C′ respectively, thereby constructing a virtual fluorescence boundary. Even if the fluorescence signal disappears or the surface of the organ is deformed and offset, the tracking algorithm based on image features can still achieve real-time tracking of the boundary, thereby achieving the effect of time-lapse imaging.

[0097] like Figure 5 As shown in the figure, the left part is the initial moment of fluorescence imaging, at which the boundary of the target area is clearly visible. Through fluorescence segmentation and boundary extraction technology, the target boundary is accurately identified, and the extracted virtual boundary is marked with a red line.

[0098] The right part of the figure is the moment when the fluorescence signal has faded, and traditional fluorescence imaging can no longer provide effective boundary information. The present invention uses a key point tracking algorithm based on image features to track the initial boundary position in real time, and the red line is the delayed virtual boundary formed by tracking. The boundary is not statically maintained, but is dynamically detected, matched and updated based on the image features of the organ surface, so that in complex dynamic scenes such as respiratory movement, instrument interference or tissue deformation, it can still accurately follow the boundary flexibly with the movement of the target area.

[0099] As mentioned above, the present invention not only realizes the time-delay display of the fluorescence boundary, but also has excellent dynamic adaptability, providing a continuous and stable visual reference for intraoperative navigation.

[0100] like Figure 4 As shown, the fluorescence imaging image is collected from the laparoscope image processor, the processing unit processes the fluorescence imaging image, and finally, the display screen displays the processing result. A dedicated image processing chip (such as FPGA) is used for boundary detection and tracking to ensure real-time processing of large amounts of fluorescence imaging images during surgery.

[0101] Figure 6 is a block diagram showing an image processing apparatus according to some embodiments of the present disclosure. Figure 6 As shown, the image processing apparatus 600 includes an image acquisition module 610 , an image segmentation module 620 , a probability map processing module 630 , a final boundary sampling module 640 , a displacement vector determination module 650 , a final boundary recalculation module 660 , and a final boundary reconstruction module 670 .

[0102] An image acquisition module 610 is configured to acquire an ICG fluorescence imaging image; The image segmentation module 620 is configured to perform global segmentation on the ICG fluorescence imaging image using a trained UNet deep learning network to obtain a probability map of the watershed boundary; A probability map processing module 630 is configured to process the probability map to obtain a final boundary; A final boundary sampling module 640 is configured to sample the final boundary to obtain a set of key point coordinates; The displacement vector determination module 650 is configured to determine the displacement vector of the key point coordinate set within a preset time interval; The final boundary recalculation module 660 is configured to recalculate the boundary if the average confidence of the consecutive K frames is lower than the threshold. , recalculate the final boundary and refresh the key point coordinate set.

[0103] The final boundary reconstruction module 670 is configured to reconstruct the dynamic change of the final boundary according to the key point coordinate set and the displacement vector.

[0104] In the device of the disclosed embodiment, the trained UNet deep learning network is used to perform global segmentation on the ICG fluorescence imaging image to obtain the probability map of the watershed boundary, the Canny edge detection algorithm is used to detect the ICG fluorescence imaging image to obtain the detail boundary, and the fused boundary is processed with the help of the morphological operation function to obtain the final boundary, thereby improving the accuracy of boundary extraction in dynamic and low signal-to-noise ratio environments. According to the key point coordinate set and the displacement vector, the dynamic change of the final boundary is reconstructed, which can effectively deal with the situation of organ flipping and deformation, and maintain the continuous stability of the boundary prompt, that is, extend the boundary prompt time.

[0105] Figure 7 FIG. 1 is a block diagram showing an image processing apparatus according to some other embodiments of the present disclosure. Figure 7 As shown, the image processing device 700 includes a memory 710; and a processor 720 coupled to the memory 710. The memory 710 is used to store instructions for executing the corresponding embodiments of the image processing method. The processor 720 is configured to execute the image processing method in any of the embodiments of the present disclosure based on the instructions stored in the memory 710.

[0106] Figure 8 is a block diagram showing a computer system for implementing some embodiments of the present disclosure. Figure 8 As shown, the computer system 800 may be embodied in the form of a general-purpose computing device. The computer system 800 includes a memory 810, a processor 820, and a bus 830 that connects various system components.

[0107] The memory 810 may include, for example, a system memory, a non-volatile storage medium, etc. The system memory may store, for example, an operating system, an application program, a boot loader, and other programs. The system memory may include a volatile storage medium, such as a random access memory (RAM) and / or a cache memory. The non-volatile storage medium may store, for example, instructions for executing at least one corresponding embodiment of the image processing method. The non-volatile storage medium may include, but is not limited to, a disk memory, an optical memory, a flash memory, etc.

[0108] The processor 820 can be implemented by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistors, etc. Discrete hardware components. Accordingly, each module such as the image acquisition module, the image segmentation module, the probability map processing module, the final boundary sampling module, the displacement vector determination module, the final boundary recalculation module, and the final boundary reconstruction module can be implemented by a central processing unit (CPU) running instructions in a memory that execute corresponding steps, or can be implemented by a dedicated circuit that executes corresponding steps.

[0109] The bus 830 may use any of a variety of bus architectures, including, but not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, and a Peripheral Component Interconnect (PCI) bus.

[0110] The computer system 800 may also include an input / output interface 840, a network interface 850, a storage interface 860, etc. These interfaces 840, 850, 860, the memory 810, and the processor 820 may be connected via a bus 830. The input / output interface 840 may provide a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 850 may provide a connection interface for various networked devices. The storage interface 860 may provide a connection interface for external storage devices such as a floppy disk, a USB flash drive, and an SD card.

[0111] Here, various aspects of the present disclosure are described with reference to flowcharts and / or block diagrams of methods, devices, and computer program products according to embodiments of the present disclosure. It should be understood that each frame of the flowchart and / or block diagram and the combination of frames can be implemented by computer-readable program instructions.

[0112] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable device to produce a machine, so that the processor executes the instructions to produce means for implementing the functions specified in one or more blocks in the flowchart and / or block diagram.

[0113] These computer-readable program instructions may also be stored in a computer-readable memory, which cause the computer to work in a specific manner to produce an article of manufacture, including instructions for implementing the functions specified in one or more blocks in the flowchart and / or block diagram.

[0114] The present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.

[0115] In the present disclosure, the trained UNet deep learning network is used to perform global segmentation on the ICG fluorescence imaging image to obtain the probability map of the watershed boundary, and the Canny edge detection algorithm is used to detect the ICG fluorescence imaging image to obtain the detail boundary. The fused boundary is processed with the help of the morphological operation function to obtain the final boundary, which improves the accuracy of boundary extraction in dynamic and low signal-to-noise ratio environments. According to the key point coordinate set and displacement vector, the dynamic changes of the final boundary are reconstructed, which can effectively deal with the situation of organ flipping and deformation, and maintain the continuous stability of the boundary prompt, that is, extend the boundary prompt time.

[0116] So far, the image processing method and device according to the present disclosure have been described in detail. In order to avoid obscuring the concept of the present disclosure, some details known in the art are not described. Based on the above description, those skilled in the art can fully understand how to implement the technical solution disclosed here.

[0117] Although some specific embodiments of the present disclosure have been described in detail by way of examples, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. An image processing method, characterized in that: The method comprises: Acquire ICG fluorescence imaging images; Using the trained UNet deep learning network to perform global segmentation on the ICG fluorescence imaging image to obtain a probability map of the watershed boundary; Processing the probability map to obtain a final boundary; Sampling the final boundary to obtain a set of key point coordinates includes: sampling at the extreme value of curvature according to the confidence threshold Select no less than the preset number of anchor points, and the adjacent anchor points are less than or equal to , forming the key point coordinate set, where the local curvature corresponding to the anchor point is the curvature extreme value of the corresponding pixel on the curve, and the final boundary confidence is greater than or equal to , , the difference in arc lengths of adjacent anchor points along the curve is less than or equal to ; The improved TapNet-Siamese key point tracking network is used to determine the key point coordinate set at a preset time interval The method comprises: embedding features of local image blocks centered on each key point in the reference frame and the current frame through a weight-sharing two-branch convolutional subnetwork; matching the embedding vectors using cosine similarity to obtain the displacement vector; and when the matching confidence is lower than a threshold Or the cross-frame displacement is greater than , the cross-frame cross-attention network based on the TapNet structure is called to re-estimate the displacement vector, where: Indicates the displacement threshold; the TapNet-Siamese key point tracking network includes: an input layer for receiving the reference frame, the current frame, and the key point coordinates, a dual convolution-TSM-ResNet18 backbone with fully shared weights, a Siamese metric learning branch for outputting the key point embedding vector, and a The Cross-Attention module performs cross-frame feature alignment and the displacement regression head outputs the displacement vector; If the average confidence of consecutive K frames is lower than the threshold , recalculate the final boundary and refresh the key point coordinate set; The dynamic change of the final boundary is reconstructed according to the key point coordinate set and the displacement vector.

2. The image processing method according to claim 1, characterized in that: The processing of the probability map comprises: Binarizing the probability map to obtain a rough boundary; Using the Canny edge detection algorithm to detect the ICG fluorescence imaging image to obtain detail boundaries; Merging the rough boundary with the detail boundary to obtain a fused boundary; The fused boundary is processed by means of a morphological operation function to obtain a final boundary.

3. The image processing method according to claim 1, characterized in that: Before acquiring the ICG fluorescence imaging image, the method further includes: Acquire initial ICG fluorescence imaging images; Performing grayscale processing on the initial ICG fluorescence imaging image to obtain a grayscale image; Denoising the grayscale image to obtain an ICG fluorescence imaging image; Returns the ICG fluorescence imaging image.

4. The image processing method according to claim 1, characterized in that: The ICG fluorescence imaging image is globally segmented using the trained UNet deep learning network to obtain a probability map of the watershed boundary, wherein the UNet deep learning network is optimized using a cross entropy loss function, and the parameters of the UNet deep learning network are updated using an Adam optimizer.

5. The image processing method according to claim 2, characterized in that: The ICG fluorescence imaging image is globally segmented using the trained UNet deep learning network to obtain a probability map of the watershed boundary, including: According to the formula: , and get the probability map ,in, represents the trained UNet deep learning network, represents the ICG fluorescence imaging image, represents the spatial coordinates, Represents a time variable.

6. The image processing method according to claim 5, characterized in that: The binarization process of the probability map to obtain a rough boundary includes: According to the formula: , and get a rough boundary ,in, Indicates setting a threshold.

7. The image processing method according to claim 6, characterized in that: The method of detecting the ICG fluorescence imaging image using the Canny edge detection algorithm to obtain detail boundaries includes: According to the formula: , get the detail boundary .

8. The image processing method according to claim 7, characterized in that: The step of processing the fused boundary by means of a morphological operation function to obtain a final boundary includes: According to the formula: , and obtain the final boundary ,in, represents the fusion boundary, Represents morphological operation function.

9. The image processing method according to claim 8, characterized in that: The reconstructing the dynamic change of the final boundary according to the key point coordinate set and the displacement vector includes: According to the formula: , reconstructing the dynamic change of the final boundary, where Indicates the preset time interval The displacement vectors of each key point in Indicates at time The sampled key point coordinate set, Represents the key point coordinate set after dynamic changes .

10. An image processing device, characterized in that: include: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the image processing method according to any one of claims 1 to 9 based on instructions stored in the memory.

Citation Information

Patent Citations

  • Fluorescence navigation system used in tumor surgery

    CN101770141A

  • Fluorescence imaging device applying spectrum detection and testing method

    CN112414987A

  • Structured landmark detection via topology-adapting deep graph learning

    US20230245329A1

  • System and method for matching of block and slice histological samples

    US20230326025A1

  • Methods and systems for identifying lymph nodes using fluorescence imaging data

    US20250031971A1