Transforaminal endoscopic image segmentation algorithm operation assistance method and system

CN122530231APending Publication Date: 2026-08-07GENERAL HOSPITAL OF THE NORTHERN WAR ZONE OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GENERAL HOSPITAL OF THE NORTHERN WAR ZONE OF THE CHINESE PEOPLES LIBERATION ARMY
Filing Date
2026-05-19
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供椎间孔镜经皮内镜影像分割算法操作辅助方法及系统,解决了现有技术依赖术前静态影像与外部空间定位设备进行三维配准,导致在组织动态形变下配准误差累积、操作链路繁琐,以及缺乏对动态手术风险进行预测性感知的技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530231A_ABST
    Figure CN122530231A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of image recognition and computer vision surgery auxiliary, discloses an intervertebral foramen mirror percutaneous endoscope image segmentation algorithm operation auxiliary method and system, the method comprises the following steps: acquiring intraoperative monocular RGB video stream and preprocessing; image recognition is carried out by using deep neural network, and four kinds of anatomical target semantic segmentation mask are output; stable mask is generated through dense optical flow field time domain consistency constraint; after pseudo-color coding, the original image is fused to realize augmented reality display; the instrument mask is extracted by using optical flow motion consistency, the physical distance scale factor is obtained by combining the intraoperative reference object, the instrument and nerve root distance and the predicted collision time are calculated, and any value below the threshold value is early warning. The present application constructs two-dimensional self-consistent closed loop, eliminates dynamic deformation error, realizes predictive early warning, and improves surgical safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image recognition and computer vision surgical assistance technology, specifically relating to an operation assistance method and system for percutaneous endoscopic discectomy image segmentation algorithm. Background Technology

[0002] Percutaneous endoscopic discectomy (PED) has become the mainstream minimally invasive procedure for treating lumbar disc herniation due to its advantages of minimal trauma and rapid recovery. During the procedure, the surgeon inserts the endoscope into the intervertebral foramen area through a working channel. Under continuous saline irrigation, the herniated nucleus pulposus is removed based on real-time image guidance. The intervertebral foramen area has a dense and complex anatomical structure, with nerve roots, dural sacs, dorsal root ganglia, and tiny vascular plexuses overlapping and enveloping each other. Under the conditions of irrigation fluid and point light source illumination, the color and texture differences of various soft tissues are minimal. Bleeding and tissue debris further reduce the image signal-to-noise ratio, making boundary identification extremely difficult.

[0003] Current mainstream auxiliary solutions rely on preoperative images to establish an external reference system. This involves constructing a three-dimensional anatomical geometric model including bones and nerves by fusing preoperative 3D CT and 3D MRI data from the same patient. Soft tissue deformation is simulated using a finite element model, and then the instrument coordinates acquired through optical or electromagnetic tracking are registered with the deformed model to achieve spatial navigation and early warning. Patent CN121533816A discloses a precise interventional navigation system for percutaneous endoscopic discectomy in cervical and lumbar disc herniation, belonging to the field of intelligent medical and surgical navigation technology. This system includes a multimodal information acquisition module, a channel planning module, a marker tracking module, a dynamic update module, and a safety early warning module. It constructs a high-precision initial three-dimensional anatomical model by fusing preoperative 3D CT and MRI data, automatically identifies vertebral anatomical feature points, and generates a safe working channel trajectory. Combining optical-inertial hybrid tracking and drift calibration mechanisms, it achieves highly robust spatial positioning of the insertion instrument. Based on a finite element mechanical model and real-time elastic registration, it dynamically updates the navigation coordinate system and implements graded early warnings for key structures such as nerves and the dural sac based on multi-level safety thresholds. It collects preoperative CT and MRI data, integrates them to construct a three-dimensional anatomical geometric model, uses a finite element mechanical model to simulate the dynamic deformation of soft tissue, and registers the coordinates of the optically tracked instrument with the model to achieve navigation and graded early warning.

[0004] The aforementioned technical prerequisites for establishing a three-dimensional coordinate system outside the patient's body face insurmountable physical constraints. Minor adjustments in patient positioning, spinal tremors induced by respiration, and local pressure fluctuations in the irrigation fluid collectively drive continuous nonlinear spatial deformation of key anatomical structures. A structural mismatch exists in the temporal dimension between preoperative static CT data and intraoperative dynamic reality, with registration errors accumulating and amplifying as the surgery progresses. Optical or electromagnetic tracking devices require additional sensors within the confined working channel of the endoscope, increasing system size and operational complexity.

[0005] At the same time, there is a perceptual spatial separation between the measurement of three-dimensional spatial distance and the image plane on the monitor that doctors rely on visual judgment. Early warning is based on geometric calculations of an abstract coordinate system, while surgical decisions rely on visual judgment on the monitor. There is a delay in information conversion, and the core operation of the surgery still relies on experience judgment, which is a semi-blind state. Summary of the Invention

[0006] The purpose of this invention is to provide an operation assistance method and system for percutaneous endoscopic discectomy image segmentation algorithm, which solves the technical problems of existing technologies that rely on preoperative static images and external spatial positioning equipment for three-dimensional registration, resulting in the accumulation of registration errors under dynamic tissue deformation, cumbersome operation links, and lack of predictive perception of dynamic surgical risks.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: The operation assistance method for percutaneous endoscopic discectomy image segmentation algorithm includes the following steps: Step 1: Acquire a real-time monocular RGB video stream during percutaneous endoscopic discectomy, wherein the video stream has a fixed frame rate; Step 2: Preprocess each frame of the video stream to obtain a preprocessed image sequence; Step 3: Input the preprocessed image sequence into a deep neural network model and output pixel-level semantic segmentation masks for four types of anatomical structures: nerve roots, dural sac, protruding nucleus pulposus, and bony structures. Step 4: Perform temporal consistency constraint processing on the semantic segmentation mask of the current frame based on the dense optical flow field to generate a stable semantic segmentation mask; Step 5: The stabilized semantic segmentation mask is pseudo-color encoded and synthesized with the original video stream to form an augmented reality image for display. Step six: Using the dense optical flow field generated in step four, extract the surgical instrument mask based on motion consistency segmentation; obtain the real-time physical distance scaling factor based on intraoperative reference objects; using the real-time physical distance scaling factor, convert the minimum pixel Euclidean distance between the edge of the surgical instrument mask and the edge of the nerve root mask into a physical distance, and calculate the expected collision time based on the motion information of the edges of the surgical instrument mask and the nerve root mask; when the physical distance is less than a preset distance safety threshold or the expected collision time is less than a preset time safety threshold, trigger an early warning signal.

[0008] Furthermore, in step two, the preprocessing includes: Each frame of the image is subjected to contrast-limited adaptive histogram equalization, with the contrast limit threshold set between 0.01 and 0.03. The nonlocal mean filter is used for noise reduction, and the smoothing parameter is automatically increased when the signal-to-noise ratio decreases and automatically decreased when the signal-to-noise ratio recovers.

[0009] Furthermore, in step six, the calculation of the estimated collision time based on the motion information of the surgical instrument mask edge and the nerve root mask edge specifically includes: Extract the first motion vector of key pixels on the edge of the surgical instrument mask. The second motion vector of the corresponding pixel closest to the key pixel on the edge of the nerve root mask is obtained through the dense optical flow field extracted in step four. ; Calculate the relative motion vector of the key pixel relative to the corresponding pixel. ; The formula for updating the estimated collision time based on the relative motion vector is as follows: , In the formula The estimated collision time is expressed in frames. The pixel Euclidean distance from the key pixel to the corresponding pixel is given. The magnitude of the relative motion vector is expressed in pixels per frame. The angle between the direction of the relative motion vector and the direction vector pointing from the key pixel to the corresponding pixel is given. This is a positive smoothing term with a value of 0.001. When judging If no dynamic collision trend is found, the calculation of the estimated collision time is stopped. Multiply the estimated collision time in frames by the frame period of the video stream. This yields the estimated collision time in seconds. ,in , The fixed frame rate is [the specified frame rate].

[0010] Furthermore, in step three, the deep neural network model includes an encoder network, a dilated spatial convolutional pooling pyramid module, and a decoder network. The encoder network is used to extract hierarchical features, the dilated spatial convolutional pooling pyramid module is used to extract multi-scale contextual features in parallel with multiple different dilation factors, and the decoder network is used to fuse the feature maps of each layer of the encoder, upsample them, and output the semantic segmentation mask.

[0011] Furthermore, the dilated spatial convolutional pooling pyramid module contains four parallel branches: There is one 1×1 convolution branch and three 3×3 dilated convolution branches, with the dilation factors of the three dilated convolution branches set to 6, 12 and 18 respectively.

[0012] Furthermore, the decoder network sets a feature recalibration module at each hop connection fusion node. The feature recalibration module is used to perform a global feature aggregation operation on the input feature map to generate a global descriptor in the channel dimension. After the global descriptor is processed by a gating network, a weight vector in the channel dimension is output, and the weight vector is multiplied by the input feature map.

[0013] Furthermore, the deep neural network model employs a composite loss function during training. Cross-entropy loss With Dice loss The weighted combination is expressed as: , In the formula The weighting coefficients are set to a range of 0.4 to 0.6. The formula for calculating Dice loss is as follows: , In the formula For the first The predicted probability that each pixel belongs to the target category. For the first The category label of each pixel in the actual annotation.

[0014] Furthermore, in step four, the temporal consistency constraint processing specifically includes: Maintain a circular buffer queue with a capacity of 5 frames to store the final output segmentation mask of the previous frame; calculate the normalized cross-correlation metric between the initial predicted segmentation mask of the current frame and the stable segmentation mask of the previous frame. When the normalized cross-correlation metric is lower than a preset correlation threshold, or when the average pixel-level spatial displacement between adjacent frames is detected by tracking feature points on the segmented edge contour through dense optical flow field tracking, a deformation compensation process based on dense optical flow field is initiated. In the deformation compensation process, a confidence map based on local gradient magnitude is constructed, and adaptive weighted fusion is performed on the deformation-aligned previous frame mask and the current frame prediction mask based on the confidence map.

[0015] Furthermore, the adaptive weighted fusion formula is as follows: , In the formula The output segmentation label probability value after fusion. For the deformed and aligned preceding frame mask at position The segmentation label probability value at the location, Predict the mask at the current frame position The segmentation label probability value at the location, The fusion weights are the masks from the preceding frames. The calculation formula is: , In the formula For position The local gradient magnitude at that point, The threshold for distinguishing gradients between edges and flat regions. A scaling factor used to control the smoothness of transitions.

[0016] Furthermore, in step five, the synthesized augmented reality image employs a transparency overlay formula: , In the formula The pixel values ​​of the composite image. These are the pixel values ​​of the segmentation mask after pseudo-color encoding. These are the pixel values ​​corresponding to the original video stream. This is the transparency coefficient, with a value ranging from 0.2 to 0.4.

[0017] Furthermore, in step five, the synthesized augmented reality image employs a dynamic visual warning mechanism: When the warning signal is triggered in step six, a two-dimensional Gaussian risk thermal field is constructed centered on the corresponding pixel closest to the key pixel on the edge of the neural root mask. The spatial probability distribution formula of the two-dimensional Gaussian risk thermal field is: , In the formula Image plane coordinates Risk warning weighting at the location The image plane coordinates of the corresponding pixel point. For the risk diffusion radius, and , This is a preset scaling constant. The estimated collision time is in seconds. It is a preset minimum positive time constant used to prevent division by zero; The two-dimensional Gaussian risk thermal field is mapped as a gradient pseudo-color layer that decays from red to yellow from the inside out, and this gradient pseudo-color layer is superimposed and composited with the original video stream for display; the smaller the expected collision time, the greater the risk diffusion radius. The larger the value, the greater the pixel coverage of the gradient pseudo-color layer in the augmented reality image.

[0018] Furthermore, in step six, the surgical instrument mask is extracted using a segmentation method based on motion consistency: Using the dense optical flow field calculated in step four, the difference between the rigid body motion mode of the surgical instrument and the non-rigid deformation motion mode of the surrounding soft tissue is analyzed. By classifying the motion mode of the dense optical flow field, the pixel region that conforms to the rigid body motion mode is determined as the surgical instrument mask.

[0019] Furthermore, in step six, the calculation of the estimated collision time is based on the motion vector of the key pixel, and the formula for calculating the estimated collision time in frames is as follows: In the formula The estimated collision time is expressed in frames. The Euclidean distance is the pixel distance from a key pixel on the edge of the surgical instrument mask to the nearest pixel on the edge of the nerve root mask. The motion vector magnitude of the key pixel, in pixels per frame. The angle between the direction of the key pixel's motion vector and the direction from the key pixel to the nearest point on the edge of the neural root mask. The positive smoothing term has a value of 0.001; the estimated collision time in frames is multiplied by the frame period of the video stream to obtain the estimated collision time in seconds.

[0020] Furthermore, the calculations in steps two through six are all performed in the graphics processor, and the frame data of the original video stream is directly written from the hardware buffer of the acquisition interface to the graphics processor's video memory via direct memory access.

[0021] Furthermore, the method also includes a visual quality self-diagnosis step: When the global average gradient magnitude of the input image is detected to have decreased by more than 50% compared to the initial baseline level of the surgery, or when the area of ​​the connected region of dark red pixels is determined to exceed 40% of the total image area, the overlay display of the augmented reality mask layer is automatically paused and a prompt message is displayed.

[0022] Furthermore, the time-domain consistency constraint processing further includes: When the normalized cross-correlation metric of multiple consecutive frames falls below a preset visual confidence threshold, the deformation compensation process based on the dense optical flow field is paused, and the following backup processing is performed: Before the operation, two-dimensional X-ray fluoroscopic images of the surgical segment were obtained, and the bony contour of the intervertebral foramen region was extracted from them and a two-dimensional kinematic constraint model containing only two-dimensional translation and sagittal rotation degrees of freedom was constructed. During the procedure, when the threshold condition is met, the stable segmentation mask before the current frame is used as the reference mask, and the pose change of the bony structure in the current frame is predicted using the two-dimensional kinematic constraint model. The prediction is achieved through the following process: For the current frame, a preset high-threshold Canny edge detection operator is used to extract sparse bony edge gradient points. The objective function is to minimize the chamfer distance between the bony contour in the reference mask and the bony edge gradient points extracted in the current frame. Within the degree-of-freedom space defined by the two-dimensional kinematic constraint model, the L-BFGS algorithm is used to optimize the transformation parameters to obtain the optimal two-dimensional rigid body transformation matrix. The optimal two-dimensional rigid body transformation matrix is ​​then applied to all pixel coordinates in the reference mask to generate the final output segmentation mask for the current frame.

[0023] Furthermore, in step six, the preset time safety threshold is adaptively adjusted based on the patient's real-time respiratory rate, and the dynamic threshold expression is: In the formula For time The changing dynamic time safety threshold The baseline physiological response time constant, The breathing frequency is extracted from the mean sequence of global background optical flow amplitude by Fourier transform. This is the amplitude adjustment weighting coefficient.

[0024] Furthermore, in step six, obtaining the real-time physical distance scaling factor based on intraoperative reference points specifically includes: A mark of known physical size is pre-set on the distal end region of the surgical instrument; During the procedure, the pixel region corresponding to the mark is detected on the surgical instrument mask, and the pixel width of the mark in the current frame is calculated. Divide the known physical size by the pixel width to obtain the real-time physical distance scaling factor, in millimeters per pixel; use this scaling factor to convert any pixel distance into physical distance.

[0025] In addition, this invention also discloses an operation assistance system for percutaneous endoscopic discectomy image segmentation algorithm, comprising: Endoscopic video acquisition unit, used to acquire real-time monocular RGB video stream during surgery; The edge computing unit integrates a graphics processor, which includes a preprocessing module, a deep neural network model, a temporal consistency constraint module, and a spatial conflict early warning module. The display unit is used to present augmented reality images; The human-computer interaction module is used to receive operation instructions; the preprocessing module, deep neural network model, temporal consistency constraint module and spatial conflict early warning module work together to execute the percutaneous endoscopic image segmentation algorithm operation assistance method described above.

[0026] Compared with the prior art, the present invention has the following beneficial effects: This invention eliminates the need for external spatial positioning sensors and preoperative 3D image registration, fundamentally avoiding the accumulation of registration errors caused by dynamic soft tissue deformation due to patient positioning adjustments, respiratory tremors, and fluctuations in irrigation fluid pressure, thus significantly simplifying the surgical procedure. By utilizing a real-time physical distance scaling factor based on intraoperative reference points, it for the first time maps 3D collision risk to a joint criterion of physical distance and predicted collision time within a 2D image plane, achieving predictive warnings for the dynamic approach process between instruments and nerve roots, surpassing static distance warning modes. By extracting relative motion vectors from a dense optical flow field to calculate the predicted collision time, it effectively eliminates false alarms and missed alarms caused by periodic nerve root pulsations, significantly improving warning reliability. The introduction of a backup path based on an anatomical kinematic constraint model maintains segmentation and warning functions even under extreme visual algorithm conditions such as lens contamination or blood occlusion, achieving a leap from functional degradation to continuous functional redundancy. Furthermore, a dynamic Gaussian risk thermal field transforms the abstract collision urgency into a pseudo-color warning force field of intuitive expansion and contraction at the nerve root edge, allowing the surgeon to intuitively perceive the risk situation without cognitive conversion.

[0027] Meanwhile, the end-to-end computation is completed in the graphics processor, the end-to-end latency is stably suppressed to an extremely low level, the nerve root segmentation accuracy reaches an extremely high level, the inter-frame mask flicker rate is extremely low, the distance measurement standard deviation is extremely small, the segmentation results are stable and the boundaries are clear, transforming the semi-blind operation that previously relied on the surgeon's sensory inference into a precise micro-operation with quantitative boundary guidance and proactive risk prediction, significantly improving surgical safety and operational efficiency. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0029] Figure 1 This is a flowchart of the method of the present invention.

[0030] Figure 2 This is a flowchart of the sub-process for calculating the predicted collision time based on relative motion vectors in this invention.

[0031] Figure 3 This is a flowchart of the alternative path processing sub-process of the present invention.

[0032] Figure 4 This is a flowchart of the dynamic visual warning mechanism sub-process of the present invention.

[0033] Figure 5 This is a diagram of the user interface of the system described in this invention. Detailed Implementation

[0034] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0035] The following is in conjunction with the appendix Figures 1-5 The embodiments of the present invention will be described in detail below.

[0036] Example 1: This example discloses an operation assistance system for percutaneous endoscopic discectomy image segmentation algorithm, including: Endoscopic video acquisition unit, used to acquire real-time monocular RGB video stream during surgery; The edge computing unit integrates a graphics processor, which includes a preprocessing module, a deep neural network model, a temporal consistency constraint module, and a spatial conflict early warning module. The display unit is used to present augmented reality images; The human-computer interaction module is used to receive operation instructions; the preprocessing module, deep neural network model, temporal consistency constraint module and spatial conflict early warning module work together to execute the percutaneous endoscopic image segmentation algorithm operation assistance method described below.

[0037] Throughout the entire surgical cycle, the system uses a monocular RGB video stream from the percutaneous endoscopic discectomy (PED) as the sole source data input. Through a composite computational pipeline integrating deep learning inference, spatiotemporal constraint computation, and image planar quantization analysis, it completes pixel-level real-time semantic segmentation of four types of anatomical targets: nerve roots, dural sac, herniated nucleus pulposus, and bony structures. The same segmentation mask simultaneously drives augmented reality display and spatiotemporal dual-condition safety alerts.

[0038] Step 1: Acquire real-time monocular RGB video stream during percutaneous endoscopic discectomy.

[0039] The endoscopic video acquisition unit is equipped with a medical-grade high dynamic range CMOS image sensor with a sampling depth of no less than 10 bits to meet the requirements for capturing tissue tone information in point-source irrigation fluid environments. The acquisition resolution is constant at 1920 pixels × 1080 pixels, and the frame rate is locked at 60 frames per second via a hardware clock, i.e., the frame period is... Seconds. The video stream is transmitted in real time via a high-bandwidth serial digital interface in uncompressed raw Bayer array format. The raw frame data uses a direct memory access mechanism, bypassing the central processing unit, and is written directly from the hardware buffer of the acquisition interface to a dedicated buffer in the video memory of the graphics processor embedded in the edge computing unit. This buffer uses a dual-circular queue ping-pong management mechanism, and read and write status exchanges are performed through hardware atomic instructions, with input stage latency approaching zero.

[0040] Step 2: Perform preprocessing on each frame of the video stream in the preprocessing module to obtain a high-quality preprocessed image sequence.

[0041] The preprocessing module comprises two sub-modules that operate sequentially: a contrast-limited adaptive histogram equalization sub-module and a non-local mean filtering denoising sub-module. The contrast-limited adaptive histogram equalization sub-module strictly sets the contrast limit threshold between 0.01 and 0.03. It performs histogram statistics on each frame of image using local blocks of 8 pixels × 8 pixels or 16 pixels × 16 pixels, and truncates and redistributes the grayscale values ​​exceeding the threshold. This threshold range was determined through extensive in vitro experiments based on the diffuse reflection physical model of the underwater point source for percutaneous endoscopic discectomy. This effectively suppresses the oversaturation of specular highlights on the surface of moist tissue while enhancing the texture contrast of the dark areas in the deep intervertebral foraminal space.

[0042] The smoothing parameter of the nonlocal mean filtering denoising submodule is inversely proportional to the real-time monitored image signal-to-noise ratio (SNR), which is evaluated online using the Laplacian operator response variance. When a sudden drop in SNR due to bleeding is detected, the smoothing parameter automatically increases nonlinearly to enhance denoising in flat areas; when the field of vision becomes clear again and the SNR recovers after irrigation fluid flushing, the smoothing parameter automatically decreases to maintain the sharpness of tissue edges.

[0043] The adaptive adjustment function is ,in The smoothing parameters for the current frame. The baseline smoothing parameter is set to 10. This is a real-time signal-to-noise ratio estimate, in dB. To adjust the intensity, a value of 2.0 is used. The constant used to control the attenuation rate is set to 5.0. When the signal-to-noise ratio drops from 20dB to below 10dB, the smoothing parameter can be automatically increased from 10 to approximately 25, significantly enhancing the noise reduction effect.

[0044] Step three involves inputting the preprocessed image sequence into a deep neural network model based on dual-path feature fusion to perform pixel-level semantic segmentation and output a four-channel probability map representing nerve roots, dural sac, protruding nucleus pulposus, and bony structures. This deep neural network model consists of an encoder network, a dilated spatial convolutional pooling pyramid module, and a decoder network.

[0045] The encoder network is a residual network architecture with five downsampling convolutional stages.

[0046] The first stage uses a 7×7 convolution kernel with a stride of 2 to quickly compress the input image resolution; the subsequent four stages are composed of multiple residual blocks stacked together, and the number of channels in the output feature maps of each stage are 64, 128, 256 and 512 respectively.

[0047] Each residual block contains two sets of 3×3 convolutions, batch normalization, and ReLU activation sequence operations, and residual connections are formed through element-wise addition of the identity mapping. The dilated spatial convolutional pooling pyramid module is located at the end of the encoder, extracting multi-scale contextual features in parallel from the high-level semantic feature map output by the encoder. This module contains four parallel feature computation branches: A 1×1 convolutional branch preserves the original feature resolution; The dilation factors of the three 3×3 dilated convolution branches were set to 6, 12 and 18, respectively.

[0048] Factor 6 is used to accurately capture the pixel span of small nerve root terminals, factor 12 is adapted to the boundaries of medium-scale targets such as the dural sac and herniated nucleus pulposus, and factor 18 is geared towards macroscopic structures such as the bony contour of the intervertebral foramen. The output feature maps of the four branches are concatenated along the channel dimension and then compressed by 1×1 convolution to obtain an output feature map that integrates multi-scale contextual information.

[0049] The decoder network, through a skip connection mechanism, fuses fine-grained feature maps from corresponding encoder layers stage by stage and performs bilinear interpolation upsampling. After passing through a Softmax activation function, it outputs a four-channel pixel-level posterior probability segmentation map. The composite loss function used during training of this deep neural network model is: ; In the formula, For composite loss function, For cross-entropy loss, For Dice's loss, The balancing weights range from 0.4 to 0.6. The formula for calculating Dice loss is: ; In the formula, For the first The predicted probability that each pixel belongs to the target category. For the first The class labels of each pixel in the actual annotation are summed and iterated through all pixels. In this embodiment... It is set to 0.5 to balance the contributions of the two losses.

[0050] Step 4: Perform temporal consistency constraint processing on the semantic segmentation mask of the current frame based on the dense optical flow field to generate a stable semantic segmentation mask.

[0051] Step four is used to suppress the flickering of mask edges between adjacent frames caused by endoscopic micromovements or irrigation fluid flow. A circular buffer queue with a capacity of 5 frames is maintained in the graphics processor memory to store the final output segmentation mask of the preceding frame.

[0052] For each newly predicted mask in a frame, first calculate the normalized cross-correlation metric between it and the most recent stable segmentation mask in the queue. , , in , The current frame prediction mask and the previous frame stable mask are located at the following positions: The label probability value at that location. , This represents the mean of the corresponding mask.

[0053] Simultaneously, a dense optical flow field is used to track equally spaced sampling feature points on the edges of two frames of masks, and the average spatial displacement of all tracked points is calculated. .

[0054] when Below the preset correlation threshold of 0.9, or When the flicker exceeds 3 pixels, edge flicker is detected, and a deformation compensation process based on dense optical flow field is initiated. This dual-condition judgment mechanism can avoid false triggering of compensation due to the decrease in normalized cross-correlation caused solely by changes in illumination, and also makes up for the deficiency that normalized cross-correlation cannot directly reflect local displacement.

[0055] The deformation compensation process uses the Farneback algorithm to calculate the dense optical flow field between the previous frame and the current frame, and performs pixel-level deformation alignment operation on the previous frame mask in the buffer queue to obtain the deformation-aligned previous frame mask.

[0056] A confidence map based on local gradient magnitudes is constructed to guide fusion, and the gradient magnitude of the current frame image is calculated using the Sobel operator. In regions with large gradient magnitudes, such as bone edges, the mask from the previous frame is assigned higher fusion weights to ensure edge continuity; in flat regions, such as the nucleus pulposus, where gradient magnitudes are smaller, the predicted mask from the current frame is assigned higher weights to more sensitively reflect the actual tissue displacement. The formula for adaptive weighted fusion is expressed as: ; In the formula, The output segmentation label probability value after fusion. For the deformed and aligned preceding frame mask at position The segmentation label probability value at the location, Predict the mask at the current frame position The segmentation label probability value at the location, The fusion weights of the preceding frame mask are calculated using the following formula: ; In the formula, For position The local gradient magnitude at that point, In this embodiment, a gradient distinction threshold is used to differentiate between edges and flat regions. Set to 50. The scaling factor, used to control the smoothness of the transition, is set to 0.1. After fusion, the generated stable segmentation mask is stored in a circular buffer queue.

[0057] Step 5: The stabilized semantic segmentation mask is pseudo-color encoded, synthesized with the original video stream to form an augmented reality image, and then pushed to the display unit. Different semi-transparent pseudo-colors are assigned to the four types of anatomical structures, and alpha mixing technology is applied to perform pixel-level synthesis based on the transparency superposition formula. ; In the formula, The pixel values ​​of the composite image. These are the pixel values ​​of the segmentation mask after pseudo-color encoding. These are the pixel values ​​corresponding to the original video stream. This is the transparency coefficient, ranging from 0.2 to 0.4. This range provides clear anatomical warning information while avoiding excessive obscuring of underlying original tissue details. The synthesized image is then displayed after anti-distortion mapping using the camera intrinsic parameter matrix built into the display unit.

[0058] Step six: In the two-dimensional image plane, based on the stable semantic segmentation mask and the surgical instrument mask extracted using the dense optical flow field in step four, perform a dual-condition joint conflict warning based on spatial distance and motion trend.

[0059] To achieve quantifiable distance warning in physical space, this step first obtains the real-time physical distance scaling factor based on intraoperative reference objects.

[0060] In practice: the distal region of the surgical instrument (such as the nucleus pulposus forceps) is pre-marked with an annular groove of known physical size, with an actual width of 2.0 mm.

[0061] During the procedure, on the extracted surgical instrument mask, the pixel region corresponding to the marker was detected by morphological template matching, and its pixel width was calculated. Real-time physical distance scaling factor (Unit: mm / pixel) Calculated by the following formula: The calibration process is dynamically updated every frame or every few frames to compensate for changes in the endoscopic object distance.

[0062] The extraction of surgical instrument masks utilizes the inherent physical differences in the movement patterns of surgical instruments and soft tissues: the former exhibits rigid translation, while the latter exhibits non-rigid peristalsis.

[0063] The system classifies the motion patterns of the dense optical flow field calculated in step four, analyzes the direction entropy and amplitude variance of the motion vector in the neighborhood of each pixel, and identifies the connected regions of pixels that conform to rigid body motion characteristics as surgical instrument masks. In the image plane, the Euclidean distance from each pixel on the edge of the surgical instrument mask to the edge of the nerve root mask is calculated, and the global minimum value is taken as the current minimum pixel distance. To achieve real-time computation, the neural root mask is pre-transformed using Euclidean distance, ensuring that the query operation for each pixel has a constant time complexity. Current minimum physical distance. The preset distance safety threshold is set to 2.0mm corresponding to physical space. When the value is less than 2.0 mm, the first type of warning is triggered.

[0064] The system analyzes the motion information of the surgical instrument mask edge and the nerve root mask edge in parallel, and calculates the estimated collision time of the surgical instrument approaching the nerve root.

[0065] The key pixel closest to the nerve root on the edge of the surgical instrument mask is located, and its motion vector is obtained through optical flow. The angle between this motion vector and the direction vector pointing towards the nearest point on the nerve root edge is calculated. The estimated collision time is calculated only when the angle is within the range of -90 degrees to 90 degrees, indicating a confirmed fatal component of instrument movement towards the nerve root. The formula for the estimated collision time in frames is: ; In the formula, The estimated collision time is expressed in frames. The Euclidean distance from the key pixel to the nearest pixel on the edge of the neural root mask is given. The motion vector magnitude of the key pixel, in pixels per frame. The angle between the direction of the key pixel's motion vector and the direction from the key pixel to the nearest point on the edge of the neural root mask. This is a positive smoothing term with a value of 0.001, used to prevent division by zero anomalies.

[0066] Convert time in frames to seconds. ,in Second.

[0067] The system presets a time safety threshold based on human neuromuscular reaction time, with a baseline value set at 0.3 seconds. When The safe distance threshold has not yet been reached, but At a certain time, a second type of preparatory warning signal is triggered, prompting the operator to slow down or adjust the instrument path through visual and auditory means.

[0068] Furthermore, to improve the accuracy of early warning in dynamic pulsating tissue environments, a more preferable method for calculating the predicted collision time is to use relative motion vectors. After locating key pixels on the edge of the surgical instrument mask, the motion vector of the nearest corresponding pixel on the edge of the nerve root mask is simultaneously acquired via optical flow field. Calculate the relative motion vector And based on this, calculate the time in units of frames. Finally, it is converted to seconds. This method eliminates false alarms or missed alarms caused by the pulsation of the nerve root itself from a physical mechanism perspective.

[0069] This dual-condition early warning mechanism makes a comprehensive judgment from two dimensions: static margin and dynamic trend, providing in-depth defense for surgical safety.

[0070] To ensure end-to-end real-time performance, all calculations in steps two through six are performed within the graphics processor. The frame data of the raw video stream is written directly from the acquisition interface hardware buffer to the graphics processor's video memory via direct memory access, without the need for central processing unit memory transfer. The end-to-end latency across the entire link remains stable within 20ms.

[0071] Furthermore, the method also includes a visual quality self-diagnosis step.

[0072] When the preprocessing module detects that the global average gradient magnitude of the input image has decreased by more than 50% compared to the baseline level established at the beginning of the surgery, or when pixel-level color space analysis determines that the area of ​​the dark red pixel connected region dominated by the red channel exceeds 40% of the total image area, the system automatically pauses the overlay display of the augmented reality mask layer and pops up a text prompt in the display unit, suggesting that the surgeon check the status of the endoscope lens.

[0073] Comparative Example 1: A navigation scheme based on preoperative 3D modeling and intraoperative optical spatial registration is provided as a reference. The specific scheme refers to the content disclosed in the prior art CN121533816A.

[0074] Under the same test conditions as in this embodiment, a continuous irrigation environment simulating the percutaneous endoscopic intervertebral foramen approach was simulated, and a simulated respiratory rhythmic tissue displacement with an amplitude of ±3 mm was introduced for comparative evaluation. Table 1 shows the performance comparison data of Example 1 and Comparative Example 1 under dynamic deformation environment.

[0075] Table 1. Comparison of dynamic deformation environment performance between Example 1 and Comparative Example 1; Comparative data shows that the pure two-dimensional self-consistent closed-loop technology approach represented in this embodiment has achieved substantial technological progress compared to traditional three-dimensional registration schemes in addressing dynamic tissue deformation and providing real-time accurate anatomical identification and early warning functions. The achievement of a distance measurement standard deviation of 0.08 mm relies precisely on the aforementioned dynamic calibration method for real-time physical distance scaling factors based on instrument markers, accurately mapping pixel distances to physical space.

[0076] Example 2: Based on Example 1, Example 2 further limits the skip connection fusion method of the deep neural network model decoder network in step three.

[0077] The encoder network, the hollow spatial convolutional pooling pyramid module, and steps one, two, four, five, and six of Example 2 are the same as those of Example 1.

[0078] In the standard skip connection of Example 1, the original feature maps output by each stage of the encoder are directly concatenated with the feature maps of the corresponding stages of the decoder. However, the complex imaging environment of percutaneous endoscopic irrigation fluid can cause some channels of the encoder feature map to be highly activated by noise modes such as reflection from suspended particles in the water or tissue debris. Directly participating in the concatenation of these noisy feature channels will reduce the signal-to-noise ratio of the decoder's fused input.

[0079] Example 2 inserts a feature recalibration module at the beginning of each skip connection fusion node of the decoder.

[0080] When the output feature map of a certain stage of the encoder is fed into the decoder, this module performs a two-dimensional spatial global average pooling operation on the input feature map, compressing it into a channel-dimensional global descriptor vector, thus condensing the global response energy of each feature channel. This descriptor vector is then fed into a gating network consisting of two cascaded fully connected layers: The first fully connected layer compresses the vector dimension to a preset ratio (e.g., a compression ratio of 8) and applies the ReLU activation function; the second fully connected layer restores the dimension to the original number of channels and uses the Sigmoid activation function to normalize the output value to the [0,1] interval, resulting in a weight vector with one channel dimension.

[0081] Through channel-level multiplication, the weight vector selectively enhances or suppresses each channel of the original input feature map. High-value channels with clear anatomical edges are given higher weights and enhanced, while noise-dominated channels are suppressed. The high signal-to-noise ratio feature map after feature recalibration is then fed into the subsequent stitching and upsampling stages. Example 2 achieves adaptive filtering along the channel dimension without significantly increasing the number of network parameters, improving the boundary clarity and segmentation accuracy of the final segmentation mask in the irrigating fluid environment.

[0082] Example 3: Based on Example 1, Example 3 provides two optional optimization schemes, namely the alternative path switching logic for extreme working conditions and the adaptive early warning threshold adjustment logic for physiological states.

[0083] In specific implementation, the extreme condition failure problem of temporal consistency constraint processing in step four is addressed. When lens fogging or large areas of blood stains cause severe degradation of visual information, the dense optical flow field will be filled with noise due to the complete loss of spatiotemporal matching of pixels. Mask deformation compensation relying on optical flow will not only fail to improve stability, but may also introduce greater prediction errors.

[0084] Example 3 constructs a redundant system consisting of a visually guided main path and a backup path driven by an anatomical kinematic constraint model.

[0085] Before the operation, a two-dimensional X-ray fluoroscopic image of the surgical segment was acquired using a C-arm X-ray machine. The preprocessing module extracted the clear bony outline of the intervertebral foramen from the image and constructed a two-dimensional kinematic constraint model containing only two-dimensional translation and sagittal rotation degrees of freedom based on the rigid body transformation matrix.

[0086] During the procedure, when the moving average of the normalized cross-correlation metric across multiple consecutive frames is lower than the preset visual confidence threshold of 0.6, the system determines that the visual input is unreliable, suspends the main path, and activates the backup path.

[0087] The alternative path abandons reliance on the optical flow field and instead employs a high-threshold (high hysteresis threshold set to 150) Canny edge detection operator to extract sparse but reliable bony edge gradient points from the current frame; using the stable segmentation mask of the last frame before visual degradation as the baseline mask, the set of bony contour points in it is denoted as... The set of bony edge gradient points extracted in the current frame is denoted as Within the degree-of-freedom space defined by a two-dimensional kinematic constraint model, using the chamfered distance... Given the objective function, the L-BFGS optimization algorithm is used to find the two-dimensional rigid body transformation matrix that minimizes the chamfer distance. ;Will The final output segmentation mask for the current frame is generated by applying all pixel coordinates to the reference mask. Once the visual input quality is restored, the system smoothly switches back to the main path.

[0088] As an optional implementation, the fixed time safety threshold in step six of Example 1 is replaced with an adaptive adjustment mechanism based on the patient's respiratory rate.

[0089] During local anesthesia surgery, the patient's breathing causes periodic movement of the tissues in the surgical area. The system performs a Fast Fourier Transform on the mean amplitude sequence of the global background optical flow to extract the principal spectral components representing the patient's real-time respiratory frequency. Based on this, a dynamic time safety threshold is constructed: ; In the formula, For time The dynamic time safety threshold, in seconds, is a variable. The baseline physiological reaction time constant is set to 0.3 seconds. The respiratory rate is extracted via Fourier transform, typically ranging from 0.2 Hz to 0.5 Hz. This is the amplitude modulation weighting coefficient, with a value of 0.1 seconds.

[0090] During the phase of most intense respiratory movement, the sine function value reaches its peak, and the dynamic time safety threshold is temporarily widened to improve warning sensitivity; during the plateau phase of respiratory interval, the sine function value drops, and the dynamic time safety threshold contracts, allowing the operator to perform more precise operations without triggering false alarms.

[0091] Example 4: This example overcomes the long-standing technical bias in existing surgical navigation technology of treating the soft tissue of the target area within a very short time window as a static rigid body, and solves the problem of false alarms and missed alarms caused by the periodic involuntary movement of nerve roots due to the patient's cardiopulmonary rhythm in the context of continuous saline irrigation under percutaneous endoscopic discectomy.

[0092] In minimally invasive surgery, nerve roots are enveloped by cerebrospinal fluid and tissue fluid, and they undergo periodic displacements of up to several millimeters within the physical space in response to arterial pulsation. Relying solely on the absolute motion vectors of surgical instruments to calculate the estimated collision time is highly susceptible to fatal missed detections when the instruments are suspended and the nerve root expands with the pulse, or to generating unnecessary intervention alarms when both instruments retract in the same direction, severely disrupting the surgical rhythm.

[0093] To extract relative kinematic features, the system reuses the dense optical flow field data used for temporal consistency constraints in step four, avoiding the need to develop new neural network computation branches. After locating key pixels on the edge of the surgical instrument mask, the graphics processor within the edge computing unit uses ultra-fast parallel search instructions to lock onto the nearest corresponding pixel on the edge of the nerve root mask.

[0094] The system directly reads the first motion vector of the key pixel from the optical flow field matrix residing in the video memory in half-precision floating-point format. and the second motion vector of the corresponding pixel. To eliminate local water flow disturbance noise, the system takes the median of the motion vectors in the 3×3 neighborhood around the corresponding pixel as the final second motion vector.

[0095] The arithmetic logic unit of the graphics processor performs hardware-level vector subtraction on two sets of two-dimensional vectors to obtain a relative motion vector that accurately represents the dynamic approximation trend of the software and hardware structures. Based on this relative motion state, the collision time is expected to be recalculated in the graphics processor's shader kernel: First, calculate the estimated collision time in frames. In the formula The estimated collision time is expressed in frames. This represents the pixel Euclidean distance from the key pixel to the corresponding pixel. The L2 norm of the relative motion vector, expressed in pixels per frame. It is the angle between the direction of the relative motion vector and the spatial vector pointing from the key pixel to the corresponding pixel. A value of 0.001 is used to prevent division by zero anomalies. The computation pipeline tracks the cosine term in real time. The polarity of the value indicates that the device and the nerve root are on opposite tracks on the visual projection plane. The system immediately blocks the division branch and cancels the warning assessment at the nanosecond level.

[0096] Convert time in frames to estimated collision time in seconds. .

[0097] The logic was validated using an ex vivo spinal model testing platform incorporating a multidimensional soft tissue deformation generator. A programmable high-frequency micro-infusion pump was used to inject periodic pulsatile displacements with a frequency of 1.2 Hz and a normal amplitude of 1.5 mm into the nerve root dura mater model, faithfully reproducing real physiological pulsations. Comparative test data are summarized in Table 2.

[0098] Table 2 Comparison of relative motion warning performance for nerve root pulsation scenarios; Data confirms that by using relative motion vectors as time estimation anchors, the system eliminates the spatial early warning blind spot induced by tissue creep from the underlying physical mechanism, with almost no increase in computing power burden.

[0099] Example 5: This example further expands the augmented reality presentation mechanism of early warning information by constructing a dynamic Gaussian risk thermal field directly mapped onto the original image plane, thereby overcoming the technical defects of traditional visual pop-ups or buzzer alarms that cause doctors to shift their gaze and delay cognitive decoding.

[0100] During the high-pressure operation of minimally invasive dissection, the doctor is in a state of high visual focus, and any external interruption across modalities will trigger the surgeon's accommodation spasm and operation pause.

[0101] This approach enables intuitive perception of risk states by reshaping the spatial force field at the edge of the target area. Once the time or distance dimension reaches the security defense threshold, the graphics processor's fragment shader will take over the rendering pipeline.

[0102] The rendering engine extracts the coordinates of the closest point at the edge of the neural root mask. As the source poles, a two-dimensional Gaussian risk thermal field is generated in the frame buffer. The spatial probability distribution of this virtual force field is governed by a mathematical model. Strict constraints.

[0103] To transform the abstract concept of time urgency into a visible spatial area, the system establishes a reciprocal mapping equation. In the formula The risk diffusion radius is expressed in pixels. The preset scaling constant is set to 150 pixels per second. The estimated collision time is in seconds. This is a preset minimum positive time constant, with a value of 0.01 seconds, used to prevent division by zero when the collision time approaches zero and to limit the maximum value of the risk radius. When When approaching the human body's limit reaction time of 0.3 seconds, the risk diffusion radius... Expand to approximately 500 pixels to cover areas of clinical concern.

[0104] In the color mapping stage, the probability weight matrix This is converted into a dynamic gradient pseudo-color layer. The high-probability core area is given a highly saturated red with an alpha channel opacity of 0.8, while the outer edges are smoothly decayed to a highly transparent yellow according to a Gaussian curve. This layer is then merged with the soft tissue segmentation mask output from step five, and the final alpha blending operation is performed with the original video stream.

[0105] The clinically apparent effect is as follows: when surgical instruments accelerate towards the nerve root within a narrow space, a red warning circle driven by a nonlinear division formula bursts open from the edge of the nerve root, engulfing the approaching instrument's tip with intense visual tension; when the surgeon decelerates or withdraws the instrument, the crisis is averted, and the red light rapidly contracts and dissipates towards its extreme point. This rendering mechanism reduces and solidifies high-dimensional predictive data into image features, ensuring the visual warning center accurately follows the non-rigid creep boundary of soft tissue, thus improving the closed-loop response rate of human-machine collaboration.

[0106] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0107] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An operational assistance method for percutaneous endoscopic discectomy image segmentation algorithm, characterized in that, Includes the following steps: Step 1: Acquire a real-time monocular RGB video stream during percutaneous endoscopic discectomy, wherein the video stream has a fixed frame rate; Step 2: Preprocess each frame of the video stream to obtain a preprocessed image sequence; Step 3: Input the preprocessed image sequence into a deep neural network model and output pixel-level semantic segmentation masks for four types of anatomical structures: nerve roots, dural sac, protruding nucleus pulposus, and bony structures. Step 4: Perform temporal consistency constraint processing on the semantic segmentation mask of the current frame based on the dense optical flow field to generate a stable semantic segmentation mask; Step 5: The stabilized semantic segmentation mask is pseudo-color encoded and synthesized with the original video stream to form an augmented reality image for display. Step six: Using the dense optical flow field generated in step four, extract the surgical instrument mask based on motion consistency segmentation; Obtain a real-time physical distance scaling factor based on intraoperative references; using the real-time physical distance scaling factor, convert the minimum pixel Euclidean distance between the surgical instrument mask edge and the nerve root mask edge into a physical distance, and calculate the estimated collision time based on the motion information of the surgical instrument mask edge and the nerve root mask edge; An early warning signal is triggered when the physical distance is less than a preset distance safety threshold or the expected collision time is less than a preset time safety threshold.

2. The operation assistance method for percutaneous endoscopic image segmentation algorithm according to claim 1, characterized in that, In step two, the preprocessing includes: Each frame of the image is subjected to contrast-limited adaptive histogram equalization, with the contrast limit threshold set between 0.01 and 0.

03. The nonlocal mean filter is used for noise reduction, and the smoothing parameter is automatically increased when the signal-to-noise ratio decreases and automatically decreased when the signal-to-noise ratio recovers.

3. The operation assistance method for percutaneous endoscopic image segmentation algorithm according to claim 1, characterized in that, Step six, which involves calculating the estimated collision time based on the motion information of the surgical instrument mask edge and the nerve root mask edge, specifically includes: Extract the first motion vector of key pixels on the edge of the surgical instrument mask. The second motion vector of the corresponding pixel closest to the key pixel on the edge of the nerve root mask is obtained through the dense optical flow field extracted in step four. ; Calculate the relative motion vector of the key pixel relative to the corresponding pixel. ; The formula for updating the estimated collision time based on the relative motion vector is as follows: , In the formula The estimated collision time is expressed in frames. The pixel Euclidean distance from the key pixel to the corresponding pixel is given. The magnitude of the relative motion vector is expressed in pixels per frame. The angle between the direction of the relative motion vector and the direction vector pointing from the key pixel to the corresponding pixel is given. This is a positive smoothing term with a value of 0.

001. When judging If no dynamic collision trend is found, the calculation of the estimated collision time is stopped. Multiply the estimated collision time in frames by the frame period of the video stream. This yields the estimated collision time in seconds. ,in , The fixed frame rate is [the specified frame rate].

4. The operation assistance method for percutaneous endoscopic image segmentation algorithm according to claim 1, characterized in that, In step three, the deep neural network model includes an encoder network, a dilated spatial convolutional pooling pyramid module, and a decoder network. The encoder network is used to extract hierarchical features, the dilated spatial convolutional pooling pyramid module is used to extract multi-scale contextual features in parallel with multiple different dilation factors, and the decoder network is used to fuse the feature maps of each layer of the encoder, perform upsampling, and output the semantic segmentation mask.

5. The operation assistance method for percutaneous endoscopic image segmentation algorithm according to claim 4, characterized in that, The dilated spatial convolutional pooling pyramid module contains four parallel branches: There is one 1×1 convolution branch and three 3×3 dilated convolution branches, with the dilation factors of the three dilated convolution branches set to 6, 12 and 18 respectively.

6. The operation assistance method for percutaneous endoscopic image segmentation algorithm according to claim 4, characterized in that, The decoder network sets up a feature recalibration module at each hop connection fusion node. The feature recalibration module is used to perform a global feature aggregation operation on the input feature map to generate a global descriptor in the channel dimension. After the global descriptor is processed by a gating network, a weight vector in the channel dimension is output, and the weight vector is multiplied by the input feature map.

7. The operation assistance method for percutaneous endoscopic image segmentation algorithm according to claim 4, characterized in that, The deep neural network model uses a composite loss function during training. Cross-entropy loss With Dice loss The weighted combination is expressed as: , In the formula The weighting coefficient is set to a value between 0.4 and 0.

6. The formula for calculating Dice loss is as follows: , In the formula For the first The predicted probability that each pixel belongs to the target category. For the first The category label of each pixel in the actual annotation.

8. The operation assistance method for percutaneous endoscopic image segmentation algorithm according to claim 1, characterized in that, Step four, specifically the temporal consistency constraint processing, includes: Maintain a circular buffer queue with a capacity of 5 frames to store the final output segmentation mask of the previous frame; calculate the normalized cross-correlation metric between the initial predicted segmentation mask of the current frame and the stable segmentation mask of the previous frame. When the normalized cross-correlation metric is lower than a preset correlation threshold, or when the average pixel-level spatial displacement between adjacent frames is detected by tracking feature points on the segmented edge contour through dense optical flow field tracking, a deformation compensation process based on dense optical flow field is initiated. In the deformation compensation process, a confidence map based on local gradient magnitude is constructed, and adaptive weighted fusion is performed on the deformation-aligned previous frame mask and the current frame prediction mask based on the confidence map.

9. The operation assistance method for percutaneous endoscopic image segmentation algorithm according to claim 8, characterized in that, The adaptive weighted fusion formula is as follows: , In the formula The output segmentation label probability value after fusion. For the deformed and aligned preceding frame mask at position The segmentation label probability value at the location, Predict the mask at the current frame position The segmentation label probability value at the location, The fusion weights are the masks from the preceding frames. The calculation formula is: , In the formula For position The local gradient magnitude at that point, The threshold for distinguishing gradients between edges and flat regions. A scaling factor used to control the smoothness of transitions.

10. A percutaneous endoscopic discectomy image segmentation algorithm operation assistance system, characterized in that, include: Endoscopic video acquisition unit, used to acquire real-time monocular RGB video stream during surgery; The edge computing unit integrates a graphics processor, which includes a preprocessing module, a deep neural network model, a temporal consistency constraint module, and a spatial conflict early warning module. The display unit is used to present augmented reality images; The human-computer interaction module is used to receive operation instructions; the preprocessing module, the deep neural network model, the temporal consistency constraint module, and the spatial conflict early warning module work together to execute the operation assistance method of the percutaneous endoscopic image segmentation algorithm as described in any one of claims 1 to 18.

Citation Information

Patent Citations

  • Precise interventional navigation system for intervertebral foramen mirror for cervical and lumbar disc

    CN121533816A