A visual perception-based method for intraoperative gauze counting and tracking

By acquiring fluorescent fingerprint vectors and entangled Kalman filter models from multispectral image sequences and combining them with an adaptive weight adjustment mechanism, the stability and consistency issues of gauze identification and tracking in surgical procedures are resolved, thereby improving the robustness and accuracy of gauze counting and tracking.

CN121259266BActive Publication Date: 2026-03-10THE SEVENTH MEDICAL CENTER OF PLA GENERAL HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511811987.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-10
Estimated Expiration
2045-12-04

AI Technical Summary

Technical Problem

Existing technologies for identifying and tracking gauze during surgical procedures suffer from challenges such as variability in identification features, difficulties in maintaining continuity in dynamic tracking, and static and non-adaptive issues in information fusion, leading to a decline in identification reliability and tracking consistency.

Method used

Fluorescent fingerprint vectors obtained from multispectral image sequences are used as the inherent identifiers of gauze. Data fusion is performed using an entangled state Kalman filter model, and an adaptive weight adjustment mechanism based on signal quality is designed. The robustness and accuracy of gauze counting and tracking are improved through a spatiotemporal tracking model and linkage control actions.

Benefits of technology

By introducing fluorescent fingerprint vectors and entangled Kalman filter models, the stability problem of gauze identity features is solved, the accumulation of tracking errors is suppressed, and the system achieves high robustness and accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121259266B_ABST
    Figure CN121259266B_ABST
Patent Text Reader

Abstract

This invention provides a visual perception-based method for intraoperative gauze counting and tracking, belonging to the field of visual pattern recognition technology. By constructing a collaborative perception framework integrating multispectral physical fingerprint recognition, visual spatiotemporal tracking, and adaptive entangled state filtering, this invention solves the core challenges of existing technologies in target identification and persistent tracking. By introducing highly robust fluorescent fingerprints as identity anchors and utilizing an adaptive fusion algorithm to achieve deep coupling and mutual correction between physical identity and spatiotemporal trajectory, absolute identification of each target and high-precision, highly robust continuous tracking are ultimately achieved. This improves the identification accuracy in severely polluted and occluded environments, effectively suppresses error accumulation during long-term tracking, and enhances the adaptive capability and reliability of the entire system in dynamic and complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual pattern recognition technology, specifically to a method for intraoperative gauze counting and tracking based on visual perception. Background Technology

[0002] With the deep integration of sensor technology and artificial intelligence algorithms, multimodal information processing has become a key technological path to improve the reliability of target perception and state estimation in complex scenarios. In fields such as automated monitoring, intelligent interaction, and high-precision tracking, integrating data sources from different physical dimensions (including optical, spectral, and spatiotemporal) to achieve a comprehensive and accurate characterization of the target's state is a significant trend in current technological development. Especially in application scenarios with high requirements for system reliability and security, how to design a robust multi-source information fusion framework to overcome the inherent limitations of a single sensing modality has become a core focus of research in this field.

[0003] Currently, in surgical applications, existing technologies face the following main challenges in achieving continuous tracking and accurate counting of gauze:

[0004] The volatility and unreliability of identification features: Existing technologies generally rely on the visual appearance features (color, shape) of targets for identification. However, in real-world environments, these features are highly susceptible to contamination (blood staining), deformation, or adhesion, leading to a decrease in the stability and distinguishability of the features. For example, in the technical solution with publication number CN120747867A, the count is adjusted by determining whether the gauze is white or red. This method risks a significant reduction in the reliability of identification when the gauze is partially or completely covered by blood, potentially leading to counting errors.

[0005] The challenge of continuity and consistency in dynamic tracking: While vision-based spatiotemporal tracking models can maintain dynamic tracking of targets to a certain extent, they are prone to tracking loss or identity swapping when targets are occluded for extended periods, move rapidly, or become confused with multiple targets. These models lack an intrinsic, unaffected, absolute identity anchor, causing tracking errors to accumulate over time and ultimately disrupting the consistency of counting.

[0006] The static and non-adaptive nature of information fusion: Even when some solutions attempt to combine multiple types of information, their fusion strategies typically employ fixed weights or simple logical rules. This static fusion mechanism cannot adapt to the dynamic changes in the signal quality of various data sources in real-world scenarios. When the quality of a particular data source deteriorates, the system cannot intelligently reduce its trust in it, and vice versa. This limits the overall performance of the fusion system to its "weakest link," and its robustness in complex dynamic environments needs further improvement. Summary of the Invention

[0007] The purpose of this invention is to provide a visual perception-based method for intraoperative gauze counting and tracking, in order to solve the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] A visual perception-based method for intraoperative gauze counting and tracking, comprising the following steps:

[0010] S1: Acquire a multispectral image sequence and extract first data from the multispectral image sequence to characterize the inherent physical identity of at least one piece of gauze, wherein the first data is a fluorescent fingerprint vector;

[0011] S2: Acquire a visual image sequence and process the visual image sequence through a spatiotemporal tracking model to obtain second data for characterizing the dynamic motion trajectory of at least one piece of gauze, wherein the second data includes motion states characterizing the position and velocity of the gauze;

[0012] S3: Based on a preset entangled state Kalman filter model, the first data and the second data are fused. The fusion process includes: jointly estimating the motion state of the gauze and the identity confidence corresponding to the fluorescent fingerprint vector in a unified state vector, and periodically correcting the identity confidence using the first data to suppress the error accumulation of the spatiotemporal tracking model.

[0013] S4: Based on the result of the fusion process, generate a set of corrected state vectors representing the absolute identity of the gauze and its corrected motion trajectory;

[0014] S5: Based on the corrected state vector, execute a linkage control action, which includes at least: feeding back the corrected motion trajectory to the spatiotemporal tracking model to calibrate its tracking reference, and guiding the excitation light source to adjust its excitation parameters for the gauze when the identity confidence in the corrected state vector is lower than a preset control threshold.

[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: Addressing the issue of the volatility of identity features, it introduces a "fluorescent fingerprint vector" obtained based on multispectral excitation as an inherent physical identity identifier for the target. This fingerprint originates from a stable fluorescent probe within the material, possessing high resistance to contamination and deformation, providing each target with a near-absolute identity anchor point that remains unchanged regardless of the external environment, fundamentally solving the unreliability problem caused by relying on volatile apparent features.

[0016] To address the challenge of continuity in dynamic tracking, a fusion model based on an entangled Kalman filter is proposed. This model incorporates the target's motion state (position and velocity) and identity confidence into a single state vector for joint estimation. The motion trajectory provided by the visual tracking model and the identity information provided by the fluorescent fingerprint are no longer simply superimposed, but rather deeply entangled and mutually corroborating within the filter. The fluorescent fingerprint periodically and with high confidence corrects the identity information, effectively suppressing identity drift and error accumulation caused by prolonged occlusion or target confusion in pure visual tracking models.

[0017] A further "adaptive weight adjustment mechanism based on real-time evaluation of dual-source signal quality" was designed. This mechanism monitors the intensity of the fluorescence signal and the stability of visual tracking in real time, and dynamically adjusts their weights during the fusion process based on this. This enables the fusion filter to intelligently "trust" the current higher-quality data source, realizing the transformation from static fusion to context-adaptive fusion, and improving the overall robustness and accuracy of the system under various complex conditions. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the process of gauze counting and tracking.

[0019] Figure 2 This is a schematic diagram illustrating the execution logic of steps S1 to S3 of the present invention;

[0020] Figure 3 This is a schematic diagram illustrating the principle of acquiring fluorescent fingerprint vectors from the gauze of the present invention.

[0021] Figure 4 This is a schematic diagram illustrating the execution logic of steps S4 to S5 of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0023] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various elements, but unless otherwise stated, these elements are not limited by these terms. These terms are used only to distinguish one element from another.

[0024] Example 1:

[0025] Please see Figures 1 to 4 The present invention provides a technical solution:

[0026] A method for intraoperative gauze counting and tracking based on visual perception includes the following steps:

[0027] S1: Acquire a multispectral image sequence during the surgical procedure, and extract first data from the multispectral image sequence to characterize the inherent physical identity of at least one piece of gauze, wherein the first data is a fluorescent fingerprint vector;

[0028] S2: Acquire a sequence of visual images during the surgical procedure and process the sequence of visual images using a spatiotemporal tracking model to obtain second data for characterizing the dynamic motion trajectory of at least one piece of gauze, wherein the second data includes motion states characterizing the position and velocity of the gauze.

[0029] S3: Based on the preset entangled state Kalman filter model, the first data and the second data are fused. The fusion process includes: jointly estimating the motion state of the gauze and the identity confidence corresponding to the fluorescent fingerprint vector in a unified state vector, and using the first data to periodically correct the identity confidence in order to suppress the error accumulation of the spatiotemporal tracking model.

[0030] S4: Based on the results of the fusion processing, a set of corrected state vectors representing the absolute identity of the gauze and the corrected motion trajectory are generated.

[0031] S5: Based on the corrected state vector, execute the linkage control action. The linkage control action includes at least: feeding back the corrected motion trajectory to the spatiotemporal tracking model to calibrate its tracking reference, and when the identity confidence in the corrected state vector is lower than the preset control threshold, guiding the excitation light source to adjust its excitation parameters for the gauze.

[0032] It should be noted that: Figure 1 The demonstration showcases a complex real-world application scenario. Surgical lamps 1-2, integrating a multispectral light source and imaging sensor, emit sequential excitation light 3 with different spectral characteristics into the surgical field below. The target gauze 4 in the surgical field is in a non-ideal state: it is wrinkled and deformed, and simultaneously partially obscured by bloodstains 5 and surgical instruments 6. Imaging sensor 2 first captures a raw image 7 containing all visual elements; in this noisy image, the fluorescence signal is weak and incomplete. Subsequently, this raw image 7 is fed into the core signal extraction algorithm module 8. This module performs crucial image processing steps, accurately separating effective fluorescence signal points from the complex background and obstructions, and filtering out all irrelevant visual interference. The result is a clean signal image 9, which retains only the true spatial distribution information of the fluorescent probes. Finally, based on this purified signal image 9, the system constructs an accurate fluorescent fingerprint vector 10 that uniquely represents the gauze.

[0033] Step S1 represents fluorescent fingerprint identity feature extraction; Step S2 represents visual spatiotemporal motion trajectory tracking; Step S3 represents cross-modal fusion and error correction based on entangled state filtering; Step S4 represents the generation of the corrected state vector; Step S5 represents closed-loop linkage control based on the state vector.

[0034] The core technical features to be elaborated in this embodiment are: a method for obtaining high-dimensional fluorescence fingerprint vectors through multispectral sequential excitation; a method for maintaining dynamic motion trajectories using a spatiotemporal graph neural network; and an entangled state Kalman filter fusion mechanism that performs asymmetric weight updates based on data source availability.

[0035] Using fixed first and second weights cannot adapt to the real-world challenges of dynamically changing signal quality during surgery. When the fluorescence signal weakens due to shallow occlusion, or when visual tracking becomes unusually stable due to a simple scene, the fixed high and low weight strategy loses its optimality. To address this issue, this invention proposes an "adaptive weight adjustment mechanism based on real-time evaluation of dual-source signal quality." This mechanism no longer treats weights as static configurations but dynamically calculates them as functions of real-time signal quality. This allows the fusion process to intelligently and in real-time adjust the level of trust in different data sources, achieving a technological leap from "static fusion" to "context-adaptive fusion," and improving the system's robustness and accuracy.

[0036] Further explanation: In S1, the fluorescent fingerprint vector is obtained in the following way:

[0037] The gauze was sequentially excited using a multispectral excitation system, and fluorescence response images generated at different excitation wavelengths were acquired simultaneously. In this way, a fluorescence fingerprint vector characterizing the response properties of the gauze in a multidimensional spectral space was constructed.

[0038] It should be noted that the multispectral excitation system is a near-infrared (NIR) LED array integrated into the surgical shadowless lamp. The gauze is pre-cured with at least two biocompatible fluorescent probes with different excitation / emission spectra in the NIR band. During data acquisition, a Python script controls the LED array to rapidly scan and excite the surgical field according to the excitation sequence (including parameters such as wavelength, duration, and intensity) defined in an external configuration file. Simultaneously, an InGaAs camera equipped with a bandpass filter captures fluorescence response images of the corresponding bands. The image processing module then calculates the fluorescence intensity change curve of each pixel under the excitation sequence and can further calculate its fluorescence lifetime decay constant. These multidimensional data are combined into a high-dimensional, highly discriminative fluorescent fingerprint vector, which serves as the absolute physical ID of the gauze.

[0039] In this invention, gauze pre-cured with fluorescent probes is used to give each piece of gauze a stable, unique, and non-invasive spectral "identity card" that can be detected in the near-infrared (NIR) band. A preferred embodiment of its preparation method specifically includes the following steps:

[0040] Gauze base: 100% pure cotton medical degreased gauze that meets ISO13485 medical device quality management system certification. The main chemical component of this gauze base is cellulose, whose molecular chain is rich in hydroxyl (-OH) functional groups that can be chemically reacted.

[0041] Fluorescent probes: Select at least two biocompatible fluorescent probe molecules with different chemical structures and distinct excitation and emission spectral peaks in the near-infrared band (700 nm-900 nm).

[0042] In a preferred embodiment, two chemically modified cyanine-7 (Cy7) dyes are selected as fluorescent probe A and fluorescent probe B. Cyanine-7 (Cy7) and its derivatives are among the most widely used and thoroughly studied near-infrared fluorescent dyes in the field of biomedical imaging. Numerous in vitro cell experiments and in vivo animal model studies have confirmed that Cy7 dyes exhibit recognized low cytotoxicity and good biocompatibility at concentrations suitable for use as imaging contrast agents. Fluorescent probe A has an excitation peak at 750 nm and an emission peak at 773 nm. Fluorescent probe B, by introducing a sulfonic acid group into its parent structure, exhibits a redshift in its spectrum, with an excitation peak at 765 nm and an emission peak at 789 nm.

[0043] To achieve "curing" with the gauze and characterize a stable covalent bond, both probe molecules were pre-modified at their ends with N-hydroxysuccinimide ester (NHS-ester) active groups that react with the hydroxyl groups of cellulose. This ensures that each fluorescent probe molecule is firmly attached to the cellulose macromolecules of the gauze. Covalent bonds are the strongest type of chemical bond and exhibit good stability in normal physiological environments (including contact with blood, tissue fluid, and body fluids at pH 7.4), without breaking.

[0044] To improve the efficiency and uniformity of subsequent bonding reactions, the gauze substrate was pretreated. The cut gauze pieces were ultrasonically cleaned in anhydrous ethanol for 15 minutes to remove any potential lipid impurities on the surface. They were then repeatedly rinsed with high-purity deionized water and completely dried in a 60°C oven. The dried, clean gauze was then immersed in a 0.5 mol / L sodium hydroxide aqueous solution and treated at room temperature for 30 minutes. This step aims to deprotonate some of the hydroxyl groups on the cellulose surface, forming more nucleophilic alkoxides, thereby significantly enhancing its reactivity with NHS-ester groups. After treatment, the gauze was rinsed with deionized water until neutral and then completely dried again.

[0045] Reaction solution preparation: Accurately weigh and dissolve the lyophilized powders of fluorescent probe A and fluorescent probe B in phosphate buffered saline (PBS) at pH 8.5. The molar ratio of the two probes is a key parameter determining the final fluorescent fingerprint, and this ratio is preset and strictly controlled. In this embodiment, the molar ratio of fluorescent probe A to fluorescent probe B is precisely controlled at 3:7.

[0046] Bonding reaction: The activated gauze was completely immersed in the prepared fluorescent probe mixture solution. The entire reaction system was placed in a constant temperature water bath at 37 degrees Celsius and reacted for 12 hours under dark conditions with gentle shaking. During this period, the NHS-ester active groups on the probe molecules underwent a nucleophilic substitution reaction with the hydroxyl groups on the gauze cellulose to form stable ester bonds, thereby firmly "fixing" the probe molecules onto the gauze fibers.

[0047] To ensure biocompatibility and eliminate background signal interference, all unreacted probe molecules existing only through physical adsorption must be thoroughly removed. The reacted gauze was repeatedly and for extended periods rinsed sequentially in PBS buffer, 50% ethanol aqueous solution, and pure deionized water. Fluorescence spectroscopy was performed on the waste wash solution between each rinsing cycle. The rinsing process was repeated until the fluorescence signal intensity of the waste solution decreased to the same level as the background noise of the blank solvent, at which point all free probes were considered to have been completely removed. The purified fluorescent gauze was then thoroughly dried in a vacuum drying oven at 40°C. Subsequently, the dried gauze was individually aseptically sealed in a GMP-compliant cleanroom. Finally, the sealed product underwent final sterilization using ethylene oxide (ETO) gas sterilization or sterilization involving gamma ray irradiation to obtain a finished medical gauze product with a unique fluorescent fingerprint, suitable for direct use in surgical environments.

[0048] Furthermore, for Figure 3 Additional explanation: Figure 3Through three interconnected areas, the entire process from microscopic physical foundations to macroscopic system applications and data-driven results is fully demonstrated. Area A3 is the microscopic mechanism area, which shows, through a magnified view, the surface of a single gauze fiber 100, on which two fluorescent probes (represented by square and triangular icons) with pre-defined proportions and types are covalently bonded. This constitutes the physical basis for the gauze's unique identification. Area A4 is the macroscopic application scenario area, where surgical lamps 1-2, integrated with near-infrared light sources, emit excitation light 3 that illuminates the specially made gauze 4 in the surgical area. The gauze then emits a fluorescent signal 204 carrying identification information, which is captured by the system. Area A5 is the data characterization area, which displays the spectral data resolved from the fluorescent signal 204. This spectral data has two unique emission peaks (302 and 303), each corresponding to a fluorescent probe. The relative intensity ratio of the two peaks directly reflects the immobilization ratio of the two probes in Area A3, thus forming a unique "fluorescent fingerprint" that can be accurately identified by the machine.

[0049] Further explanation: In S2, the spatiotemporal tracking model is a spatiotemporal graph neural network model, which is configured as follows:

[0050] By using gauze candidate targets in different video frames as graph nodes, spatiotemporal association edges between graph nodes are constructed. Node information is propagated and aggregated on the graph structure to infer and maintain the continuity of the identity and motion trajectory of each piece of gauze in the video sequence.

[0051] It should be noted that the instance segmentation model detects gauze candidate regions in each frame of the visual image. Subsequently, a Python script constructs a spatiotemporal graph from these candidate regions across frames. The node features of the graph include appearance feature vectors extracted by the CNN and the geometric attributes of the regions. Edge weights are calculated based on spatial distance between nodes, appearance similarity, and motion consistency. The hyperparameters of the spatiotemporal graph neural network model, such as the number of layers and the type of aggregation function, are managed in an external configuration file. The model learns and infers the optimal cross-frame connection paths, i.e., the motion trajectory of each piece of gauze. The performance of this method is measured by Identity Preservation Accuracy (IPA) and Multi-Object Tracking Accuracy (MOTA).

[0052] Further explanation: In S3, the entangled state Kalman filter model performs the following during its update step:

[0053] When the first data is obtained, the matching degree between the fluorescent fingerprint vector and each record in the preset fingerprint database is calculated, and the identity confidence component in the state vector is updated with a first weight based on the matching degree; when the first data is not obtained, the identity confidence component is updated with a second weight less than the first weight based on the visual appearance features of the gauze.

[0054] It should be noted that the specific calculation logic for the matching degree is as follows: The input consists of the currently acquired fluorescent fingerprint vector, which includes an N1-dimensional floating-point vector, and a reference record from a preset fingerprint database. The reference record is an N1-dimensional floating-point vector. Each dimension of the current fluorescent fingerprint vector and the reference record vector is mapped one-to-one, and the difference in values ​​along the corresponding dimension is calculated. All these differences are then squared, and finally, all squared results are summed, and the sum is squared. This step calculates the Euclidean distance between the two N1-dimensional vectors. The distance value, representing "the smaller the distance, the better the match," is converted to a matching degree value, "the higher the matching degree, the better the match," by performing an inverse function mapping. The calculated Euclidean distance is then added to a smoothing constant 1e-6 to prevent division by zero, and its reciprocal is taken. All inverse results are normalized using the Softmax function so that the sum of the matching degrees of all reference records is 1, thus forming a probability distribution. The maximum value in this probability distribution is determined as the final matching degree, and the UID of the reference record corresponding to the maximum value is assigned to the current observation object, thereby completing the confirmation of its absolute identity.

[0055] Specifically, in a specific embodiment of the present invention, the identity recognition process relies on a preset fingerprint database and an associated matching degree calculation module. The preset fingerprint database is established before the main tracking program starts, and its configuration is used to store multiple reference records corresponding to each gauze of the object to be tracked. Each reference record contains at least a unique identifier (UID) and an N1-dimensional reference record vector, which represents the standard fluorescent fingerprint of the object acquired under controlled conditions. During real-time identity matching, the matching degree calculation module receives a currently acquired N1-dimensional fluorescent fingerprint vector representing the real-time fingerprint of the observed object. Subsequently, the module iteratively compares the current fluorescent fingerprint vector with the reference record vector of each reference record in the database.

[0056] Further explanation: Before the update steps, the following are also included:

[0057] Determine a fluorescence signal quality factor characterizing the signal quality of the first data; and determine a visual tracking stability factor characterizing the tracking stability of the second data;

[0058] In the update step, the first weight and the second weight are dynamically calculated based on the fluorescence signal quality factor and the visual tracking stability factor; the value of the first weight is positively correlated with the value of the fluorescence signal quality factor, and the value of the second weight is positively correlated with the value of the visual tracking stability factor.

[0059] The following is a detailed description of the implementation of the above content: Before the calculation process of this embodiment begins, the central control module, which is implemented by a processor running a Python program, first loads all configurable running parameters from a locally stored spreadsheet file.

[0060] The fluorescence signal quality factor, denoted as FQF, is a scalar value normalized to the [0,1] interval, used to quantitatively evaluate the clarity and reliability of the fluorescence signal detected in the current frame. The closer the value is to 1, the higher the signal quality and the stronger the reliability. In this embodiment, the fluorescence response image is acquired by an InGaAs camera equipped with a bandpass filter. The original output signal is a single-channel grayscale image, where the brightness value of a pixel is proportional to the fluorescence photon flux received at the corresponding spatial location under a specific excitation wavelength. Before calculating the FQF, the original image is processed by an image segmentation module based on a U-Net semantic segmentation network. This module is pre-trained to identify "gauze regions" and "background regions" in the image and outputs corresponding binarized masks. It should be noted that, in a preferred embodiment, the "image segmentation module" is a semantic segmentation network based on a U-Net architecture; the specific implementation of this network is as follows:

[0061] The model architecture employs a standard encoder-decoder structure. The encoder consists of four downsampling blocks cascaded together, each containing two 3×3 convolutional layers (with ReLU activation) and a 2×2 max-pooling layer. The decoder, correspondingly, consists of four upsampling blocks, each containing a deconvolutional layer, a skip connection concatenating the feature maps from the corresponding encoder layer, and two 3×3 convolutional layers. The training dataset for this network comes from an internal database containing 5000 surgical field images acquired in a simulated surgical environment. All images were annotated at the pixel level by professional medical image annotators, distinguishing between "gauze regions" and "background regions." The dataset was divided into training, validation, and test sets in an 8:1:1 ratio. During training, the Dice similarity coefficient was used as the loss function, the Adam optimizer was selected, the initial learning rate was set to 0.01%, the batch size was 8, and training was performed for 100 epochs. After training, the model weights with the best validation set performance were fixed and stored, and loaded by the image segmentation module during system runtime to perform real-time image segmentation.

[0062] The following computational model is used to determine the fluorescence signal quality factor (FQF): The signal region corresponding to the gauze and its adjacent background region in the fluorescence response image are identified. Using a "gauze region" mask, all pixel values ​​belonging to that region are extracted from the original grayscale image to form the signal pixel set. Similarly, using a "background region" mask, the background pixel set is extracted. The arithmetic mean of all pixel values ​​in the "signal pixel set" is calculated to obtain the average signal intensity. The standard deviation of all pixel values ​​in the "background pixel set" is calculated to obtain the noise standard deviation. The average signal intensity is divided by the noise standard deviation to obtain the original signal-to-noise ratio (SNR). raw For the original signal-to-noise ratio (SNR) raw Perform mapping processing; read the minimum signal-to-noise ratio (SNR) threshold from the configuration file. min and maximum signal-to-noise ratio threshold (SNR) max When the original signal-to-noise ratio (SNR) raw Less than the minimum signal-to-noise ratio threshold (SNR) min When FQF is set to 0; when the original signal-to-noise ratio (SNR) is... raw Greater than the maximum signal-to-noise ratio threshold (SNR) max At this time, FQF is set to 1. When the original signal-to-noise ratio (SNR) is... raw When the signal-to-noise ratio (SNR) falls between these two values, it is mapped to the [0,1] interval using linear interpolation, thus reducing the original SNR. raw With minimum signal-to-noise ratio threshold SNR min The difference, divided by the maximum signal-to-noise ratio threshold (SNR) max With minimum signal-to-noise ratio threshold SNR min The difference between the two values ​​is used to calculate the final fluorescence signal quality factor FQF.

[0063] The visual tracking stability factor, denoted as VTSF, is a scalar value normalized to the [0,1] interval. It is used to quantitatively evaluate the smoothness and continuity of the tracking trajectory output by the spatiotemporal graph neural network over a recent period. The closer the value is to 1, the more stable the tracking state, and the lower the probability of trajectory drift or jitter. The spatiotemporal tracking model is configured to perform the following calculation logic to determine the visual tracking stability factor VTSF:

[0064] The original input data is the coordinate sequence of the two-dimensional center point of the target gauze in the past N2 frames of images, output by the spatiotemporal graph neural network model, denoted as P(t)=[x(t),y(t)], where t is the frame index representing the discrete time step; x(t) and y(t) are the horizontal and vertical coordinates of the two-dimensional center point at frame index t. This coordinate sequence is smoothed by a one-dimensional Kalman filter to obtain the smoothed coordinate sequence P'(t). The key parameters of the "one-dimensional Kalman filter" in the smoothing process are determined as follows:

[0065] The process noise covariance (Q) physically characterizes the uncertainty of the actual motion state of the target over time. In this embodiment, considering that the gauze may be subject to random disturbances such as being touched by instruments during surgery, its preferred value is set to 0.001, with a range between 0.0005 and 0.005. This value is determined by analyzing the gauze's motion trajectory in real surgical videos and statistically analyzing the standard deviation of its acceleration changes.

[0066] The noise covariance (R) is physically defined as the measurement error of the coordinate values ​​output by the visual positioning algorithm. In this embodiment, its preferred value is set to 0.01, with a range of 0.008 to 0.02. This value is determined by calculating the variance of the measurement results through 1000 consecutive positioning measurements on a stationary gauze in a static scene.

[0067] In this embodiment, based on the smoothed coordinate sequence P'(t), the discrete velocity vector sequence V(t) = P'(t) - P'(t-1) and acceleration vector sequence A(t) = V(t) - V(t-1) are calculated using the backward difference method. For each vector in the acceleration vector sequence A(t), its Euclidean norm, i.e., the magnitude of the vector, is calculated. The arithmetic mean of the norms of all acceleration vectors within a sliding time window of length N3 is calculated to obtain the average acceleration norm A. avg Obtain the maximum permissible jitter threshold A from the parameter configuration unit. max Perform an inverse linear mapping: change the mean acceleration norm A... avg Divide by the maximum permissible jitter threshold A max Obtain the jitter ratio, then subtract it from 1.0. To prevent the result from being negative, compare the calculated result with 0.0, and take the larger of the two values ​​as the final value of VTSF.

[0068] It should be noted that the minimum signal-to-noise ratio threshold (SNR) min and maximum signal-to-noise ratio threshold (SNR) max With the maximum permissible jitter threshold A max The determination logic is as follows: In this embodiment, it is the minimum signal-to-noise ratio threshold (SNR). min and maximum signal-to-noise ratio threshold (SNR) max Determine an objective numerical limit that can distinguish between "unacceptable signal quality," "acceptable signal quality," and "saturated signal quality." Let A be the maximum permissible jitter threshold. max Determine an objective numerical boundary that can distinguish between "stable tracking state" and "unstable tracking state." This boundary is determined for the minimum signal-to-noise ratio (SNR) threshold. min and maximum signal-to-noise ratio threshold (SNR) maxCalibration experiment: Prepare multiple gauze samples with standard fluorescent probes fixed on them. Use blood substitutes of different concentrations containing different proportions of hemoglobin to homogenize the samples to different degrees, and confirm the contamination level using an optical densitometer.

[0069] Under standard operating room lighting conditions, each contaminated sample was imaged at least 50 times using a multispectral imaging system, and the raw signal-to-noise ratio (SNR) for each image was calculated. raw Three senior image analysis experts conducted double-blind annotation on all acquired fluorescence images, categorizing them into three classes: "Excellent Quality," "Acceptable Quality," and "Unusable Quality." All annotation results were then compiled. The original signal-to-noise ratio (SNR) was calculated for all samples in the "Acceptable Quality" category. raw The statistical distribution of the values ​​is used, and the 5th percentile of the distribution is taken as the minimum signal-to-noise ratio threshold (SNR). min The final calibration value. Calculate the raw signal-to-noise ratio (SNR) for all samples in the "Excellent Quality" category. raw The statistical distribution of the values ​​is used, and the 95th percentile of the distribution is taken as the maximum signal-to-noise ratio (SNR) threshold. max The final calibration value.

[0070] For the maximum permissible jitter threshold A max Calibration experiment: A validation dataset containing at least 100 video clips simulating various surgical scenarios was prepared. Three experts labeled the tracking trajectories of all gauze in the dataset, categorizing them into "smooth and stable trajectory" and "trajectory jitter / abruptness". The average acceleration norm A of all trajectories was obtained. avg value.

[0071] The mean acceleration norm A is considered for two types of trajectories: "smooth and stable trajectory" and "trajectory jittering / abrupt changes". avg Find the value and plot the ROC curve of the receiver operating characteristics. Calculate the Youden exponent for each point on the ROC curve and find the mean acceleration norm A that maximizes the Youden exponent. avg This value is determined as the maximum permissible jitter threshold A. max The final calibration value is because it represents the optimal balance point for distinguishing between the two types of trajectories.

[0072] By performing the above calibration experiments, in a preferred embodiment of the present invention, the minimum signal-to-noise ratio threshold (SNR) is determined. min The maximum signal-to-noise ratio (SNR) threshold was set to 5.0. maxThe value was determined to be 45.0. This numerical range defines the effective signal-to-noise ratio (SNR) operating range; signals below 5.0 are considered to have excessive interference and are unreliable, while signals above 45.0 are considered to have reached saturation and the highest quality level. Through the aforementioned ROC curve-based calibration experiment, in a preferred embodiment of the invention, the maximum permissible jitter threshold A is determined. max It was determined to be 2.5 (unit: pixels / frame²); this value is the critical point that strikes the best balance between sensitivity and specificity in distinguishing between stable and unstable trajectories.

[0073] The core of this embodiment lies in the adaptive dual-source confidence fusion mechanism designed in step S3. It employs a fixed first weight of 0.99 and a second weight of 0.3, essentially a static, binary decision-making logic based on "present / absent" signals. This logic cannot handle continuous changes in signal quality, leading to an inability to make optimal confidence updates in intermediate states such as weak but still identifiable fluorescence signals or extremely stable visual tracking without fluorescence signals, thus reducing the overall accuracy and robustness of the system. To address these limitations, a dynamic weight allocation mechanism based on information entropy theory is proposed. Its core idea is that the higher the quality factor of the signal source, the lower its uncertainty (entropy), and therefore it should be assigned a higher weight in the fusion decision. This mechanism is implemented by a dynamic weight calculation module, whose internal working mechanism is as follows: In each update cycle, it receives the fluorescence signal quality factor FQF and the visual tracking stability factor VTSF calculated upstream. The following calculation model is executed to determine the dynamic first weight W. fdyn Second weight W vdyn :

[0074] When the fluorescence signal is available, i.e., FQF is greater than 0, the first weight W fdyn The calculation method is as follows: FQF is multiplied by the baseline confidence coefficient α read from the configuration file. α represents the prior confidence in the fluorescence signal and is set to 0.95. Meanwhile, the second weight W... vdyn The calculation method is as follows: subtract the second weight W from 1. vdyn Then multiply by VTSF. This ensures high confidence in the fluorescence signal when the fluorescence signal quality is high, while the stability of visual tracking can still adjust the weights.

[0075] When the fluorescence signal is unavailable, i.e., FQF equals 0, the second weight W vdyn The second weight W is set to 0. vdyn The calculation method is as follows: VTSF is multiplied by the visual baseline weight β read from the configuration file. β represents the baseline confidence level of visual tracking (0.4) in the absence of fluorescence signal. The calculated first weight W is then... fdyn Second weight W vdyn The output is fed into the update module of the entangled state Kalman filter.

[0076] It should be noted that the basic trust coefficient α and the visual basic weight β are given explicit physical meanings in the mathematical model, and their optimal values ​​are then determined through computational experiments. The determination method is as follows:

[0077] α and β are considered to represent the "prior confidence" of the fluorescence and visual signals, respectively, while FQF and VTSF represent the "likelihood" under the current observation. Dynamic weights are the product of the combination of these two to guide the update of posterior confidence.

[0078] We seek a set of (α,β) numerical pairs that optimizes the overall performance of the fusion tracking system on the validation dataset. The overall performance metric, OPS, is defined as the weighted sum of the aforementioned Identity Maintenance Accuracy (IPA) and Multi-Object Tracking Accuracy (MOTA), with the weights adjustable based on the application scenario.

[0079] An optimization method based on grid search is employed. The search range for α is set to [0.8, 1.0] with a step size of 0.01; the search range for β is set to [0.2, 0.6] with a step size of 0.01. All possible (α, β) combinations are traversed. For each combination, the full-process method of this invention is run on the complete validation dataset, and its OPS is calculated. The OPS values ​​corresponding to all combinations are recorded. The (α, β) combination that maximizes the global OPS value is selected as the final optimal parameter values ​​in this embodiment. Through the above-described system-level optimization experiment based on grid search, in the preferred embodiment of this invention, the parameter combination that enables the comprehensive performance index OPS to reach its optimal value is determined as follows: the preferred value for the basic trust coefficient α is 0.95, and the preferred value for the visual basic weight β is 0.4.

[0080] When the fluorescence signal is slightly blocked, causing the FQF to decrease, the first weight W fdyn It will decrease smoothly, while the second weight W vdyn Correspondingly, improvements can more "rationally" integrate the two information streams. Conversely, when visual tracking jitters due to scene changes, causing a decrease in VTSF, even without a fluorescence signal, the second weight W... vdyn This will also reduce the impact of incorrect visual information on the system state.

[0081] In this embodiment, the complete calculation process from steps S1 to S3 is scheduled and executed by the central control module, specifically including the following steps: The central control module loads all operating parameters from an external spreadsheet file, including but not limited to SNR. max SNR min A max,α,β, and initialize the spatiotemporal graph neural network model and entangled state Kalman filter (ESKF).

[0082] The multispectral imaging system sequentially excites the surgical field and simultaneously acquires fluorescence response images. The fluorescence signal processing module receives the images, and if a valid signal is detected, it calculates the fluorescence fingerprint vector and the fluorescence signal quality factor (FQF).

[0083] It should be noted that the "serialization activation" process is implemented as follows:

[0084] The input is a pre-configured excitation sequence definition table stored in an external spreadsheet file. This table contains multiple rows, each defining an excitation event and including three columns of parameters: "excitation wavelength" (in nanometers), "excitation duration" (in milliseconds), and "excitation intensity" (normalized value, between 0 and 1). During the initialization step, the central control module reads and parses this excitation sequence definition table, loading it into memory to form an excitation event queue.

[0085] The excitation controller executes excitation events sequentially according to the queue. For each event, the controller first sends a command to the multispectral LED array to illuminate the LED corresponding to the "excitation wavelength" parameter of the current event and adjust its output power to match the "excitation intensity" parameter. The controller then maintains this excitation state for a duration defined by the "excitation duration" parameter. At the instant the "excitation duration" of the excitation event ends, the excitation controller immediately turns off the current LED and triggers a synchronization signal to the InGaAs camera, enabling it to complete one image acquisition. The controller then retrieves the next excitation event from the queue and repeats the above process until all excitation events have been executed, thus completing a full sequence of excitation and synchronous acquisition.

[0086] It should be noted that, in the preferred embodiment, the visual appearance features refer to the 512-dimensional floating-point feature vector extracted by the Re-Identification (Re-ID) convolutional neural network. This feature vector is used as a graph node attribute of the spatiotemporal graph neural network in step S2 and is used in step S3 to assist in updating the identity confidence when fluorescence signals are missing. The specific implementation of the Re-Identification network is as follows: The network's model architecture is based on a pre-trained ResNet-50 network as its backbone. The final classification layer of the original ResNet-50 is removed, and a new fully connected layer is then connected. The output dimension of this fully connected layer is set to 512, thus forming an encoder capable of mapping any input gauze image to a 512-dimensional feature space.

[0087] The training dataset for this network contains 100,000 image slices collected from over 1,000 different physical gauze instances. Each image slice is precisely labeled with its unique gauze ID. During training, a triplet-loss function is used for end-to-end optimization of the network. The working logic of this loss function is as follows: in each training batch, anchor samples, positive samples belonging to the same gauze ID as the anchor sample, and negative samples belonging to different gauze IDs are randomly selected. By optimizing the network parameters, the distance between the anchor and the positive sample is minimized, while the distance between the anchor and the negative sample is maximized in the 512-dimensional feature space, ensuring that a pre-defined boundary exists between them. The Adam optimizer is used during training, with an initial learning rate set to 0.05%. For each gauze candidate region detected by the instance segmentation model (U-Net), its image data is cropped from the original video frame and its size is normalized to 224×224 pixels. Subsequently, the normalized image is input into the trained re-identification network for forward inference, and the 512-dimensional vector output by the network is the visual appearance feature of the candidate region at this moment.

[0088] Furthermore, the core of "constructing spatiotemporal connections between graph nodes" lies in calculating the association weight between any two graph nodes (node ​​i and node j) located in different frames. The calculation logic for this weight is as follows:

[0089] The input consists of the attributes of nodes i and j, including their center coordinates and a 512-dimensional appearance feature vector extracted by a convolutional neural network. The system obtains the center coordinates of nodes i and j and calculates the Euclidean distance between them. It also obtains the appearance feature vectors of nodes i and j and calculates their cosine similarity. If the system has a previous prediction of node i's trajectory, it infers the predicted position of node i in the frame containing node j. The Euclidean distance between this predicted position and the actual position of node j is calculated as a measure of motion inconsistency. The three metrics are normalized and mapped to the range of 0 to 1. The system reads the spatial distance coefficient, appearance similarity coefficient, and motion consistency coefficient from the configuration file. Each normalized metric is multiplied by its corresponding coefficient, and the three products are summed to obtain the final association weight value. This weight value is the weight of the spatiotemporal association edge connecting nodes i and j.

[0090] The visual imaging system acquires RGB image sequences. An instance segmentation model detects candidate regions for the gauze. A spatiotemporal tracking analysis module constructs a spatiotemporal graph, and a spatiotemporal graph neural network model infers the dynamic trajectory of the gauze. Simultaneously, the visual tracking stability factor (VTSF) is calculated. The system receives the free-quantity QF (FQF) and VTSF, and calculates the dynamic first weight W in real time. fdyn and dynamic second weight W vdyn .

[0091] The update module of the entangled-state Kalman filter (ESKF) performs its core update steps. When the fluorescence signal is available, the fluorescence fingerprint vector is used with a first weight W. fdyn The identity confidence score is updated using the weights; simultaneously, visual appearance features are used as the second weight W. vdyn As a weight, it is used for auxiliary updates. When the fluorescence signal is unavailable, only the visual appearance feature is used with a second weight W. vdyn Updated as a weight.

[0092] It should be further explained that the technical feature of "updating the identity confidence component in the state vector" has the following calculation logic: at the start of each time step, the following four input data items need to be obtained;

[0093] The predicted confidence level Cpred: A priori estimate of the current identity confidence level, derived from the previous time step by the ESKF prediction step. The observed identity measurement Mid: A normalized value representing the current identity measurement result. The weights W are dynamically updated. dyn Updated weights are dynamically calculated based on the availability and quality of the current data source. Observation uncertainty R id : Preset parameters that characterize the noise level of the current identity measurement process itself.

[0094] Perform conditional checks to determine the currently available data sources, and based on this, set the "identity measurement observation value Mid" and the "dynamically updated weight W". dyn The specific value of "" is as follows: When the first data, namely the fluorescent fingerprint vector, is obtained, "the observed value Mid of identity measurement" is set to the normalized matching degree calculated above and corresponding to the best match in the database.

[0095] Dynamically update weight W dyn "Set as the first weight W" fdyn When initial data is unavailable, but visual tracking is stable: extract the visual appearance features of the currently tracked target and calculate its similarity to the historical appearance feature database. Use this similarity as the "mid observation for identity measurement." Dynamically update the weights W. dyn "Set as the second weight W" vdyn .

[0096] Subtracting the predicted confidence level Cpred from the observed value Mid of the identity measurement yields the difference, which is the "confidence level residual." This residual characterizes the deviation between the actual measurement result and the system prediction.

[0097] In this embodiment, "dynamically update weight W" dyn "It is directly used as the correction gain, and its value already includes considerations for measurement quality. The 'confidence residual' and 'dynamically updated weight W' are combined..." dyn Multiplying these values ​​yields the "confidence correction amount." The magnitude and direction of this correction amount indicate the degree of adjustment to the predicted value in this update. Adding the "confidence correction amount" to the "predicted confidence value Cpred" yields a temporary "updated confidence level." A boundary constraint is then applied to the temporary "updated confidence level": if its value is greater than 1.0, it is forcibly set to 1.0; if its value is less than 0.0, it is forcibly set to 0.0.

[0098] The value after boundary constraint processing is the final, updated identity confidence component. This value will be written into the ESKF state vector as the initial value for the next time step prediction. The ESKF output is a state vector after adaptive fusion correction, which is then used to update the gauze count and position displayed on the system interface.

[0099] The identity confidence component is a scalar value normalized to the interval [0,1]. Its dynamic change directly reflects the system's certainty in identifying a specific piece of gauze. A higher identity confidence component indicates a higher level of certainty in the entangled Kalman filter model's estimation of the current gauze's identity. Specifically:

[0100] As the identity confidence component approaches 1, it indicates that the entangled Kalman filter model has achieved a higher level of certainty in estimating the identity of the current gauze. This state is achieved through a high-weighted correction step after receiving continuous, high-quality fluorescent fingerprint vectors. A confidence value approaching 1 indicates that the gauze's identity is strongly bound to a unique identifier in the database, making subsequent tracking errors less likely to cause identity confusion.

[0101] When the value of the identity confidence component approaches 0, it indicates that the entangled Kalman filter model is in a highly uncertain or ambiguous state regarding the identity estimation of the current gauze. This state occurs because no effective fluorescence signal can be obtained for a long time, and visual tracking is also unstable. As a result, historical confidence information is continuously attenuated due to process noise in the filter's prediction-update cycle, ultimately leading to the loss of effective identity information.

[0102] The dynamic evolution of the identity confidence component is primarily controlled by two key real-time input parameters: the fluorescence signal quality factor (FQF) and the visual tracking stability factor (VTSF). These two factors, through a nonlinear dynamic weight allocation mechanism, jointly determine the magnitude and direction of each state update.

[0103] The fluorescence signal quality factor (FQF) exhibits a direct, dominant, and strong positive correlation with the update gain of the identity confidence component. Under otherwise constant conditions, an increase in FQF leads to a non-linearly faster convergence of the identity confidence component towards 1.0. Conversely, a decrease in FQF slows down this convergence process. When the fluorescence signal is acquired, the first weight W... fdyn The calculation method is the product of FQF and the basic confidence coefficient α. Since α is a constant close to 1, the first weight W fdyn The value is linearly determined by FQF.

[0104] When a high-quality fluorescence signal is captured, causing the FQF to approach 1, the first weight W fdyn It also approaches α. In the ESKF update step, this high weight enables the measurement update based on the fluorescent fingerprint matching result to perform a large-scale, targeted correction on the identity confidence component, making it quickly approach 1.0.

[0105] When the fluorescence signal weakens due to factors such as occlusion, causing the FQF to decrease, the first weight W fdyn This also decreases accordingly. At this point, the weight of the measurement update becomes smaller, and the correction effect on the identity confidence component weakens.

[0106] This design reflects the logical principle that "the higher the quality of the input information, the greater its influence in the fusion decision-making process."

[0107] The value of VTSF has a conditional and auxiliary positive correlation with the attenuation suppression ability of the identity confidence component. Its main effect is evident in scenarios where fluorescence signals are missing. When FQF is 0, increasing VTSF significantly slows down the natural attenuation rate of the identity confidence component due to lack of correction, effectively "maintaining" the existing confidence level. When FQF is greater than 0, the influence of VTSF is suppressed by the FQF-dominated weighting and becomes secondary. When fluorescence signals are missing (FQF=0): the second weight W... vdyn The calculation method is the product of VTSF and the visual baseline weight (β). β is a preset constant less than α. In this case, the value of VTSF directly determines the second weight W. vdyn The size of the second weight W. When visual tracking stabilizes and VTSF approaches 1, the second weight W... vdynApproaching β. ESKF uses this weight to make a small update to the identity confidence based on the visual appearance features of the gauze; it can effectively offset the confidence decay caused by model uncertainty in the prediction step.

[0108] When visual tracking is unstable (VTSF approaches 0): Second weight W vdyn As the weight of visual updates approaches zero, the identity confidence level rapidly decays due to the uncertainty of the internal model in the absence of external correction. This design demonstrates that visual tracking itself cannot provide absolute identity information, but stable visual tracking can provide strong evidence. This mechanism improves robustness in situations involving temporary signal loss due to brief occlusion of the device.

[0109] This invention proposes a "quantitative assessment based on risk potential energy accumulation and a priority-driven active detection closed-loop mechanism." This mechanism represents the risk indicator in S4 as a continuously changing "risk potential score (RPS)," which dynamically accumulates and decays, more realistically reflecting the evolution of risk. Simultaneously, the linkage control in S5 is a "priority scheduling mechanism," which integrates the target's identity uncertainty, the changing trend of uncertainty, and the risk potential score calculated in S4. It calculates a detection priority score (PPS) for each target to be confirmed, thereby guiding the excitation light source to address the most critical issues in the optimal sequence. This improvement achieves a complete intelligent closed loop from "event-driven" to "state assessment," and then from "state assessment" to "resource optimization."

[0110] Further explanation: In S4, the corrected state vector also includes a risk indicator, which characterizes the residual risk level of the gauze. The risk indicator is determined based on the analysis of the corrected motion trajectory of the gauze. The analysis includes: detecting whether the gauze remains stationary for more than a first time threshold within a preset non-discard area, or whether its tracking trajectory is interrupted within the human body suture area. In a preferred embodiment of the present invention, the first time threshold, the coordinate range of the non-discard area, and the suture area are all defined in an external configuration file.

[0111] Further explanation: The risk indicator is the risk potential score; the determination of the risk potential score includes:

[0112] Based on discrete events such as whether the gauze enters a preset risk area or whether its tracking trajectory is interrupted, a time-varying risk baseline value is determined; and...

[0113] The continuous risk accumulation term is determined based on the degree of risk in the vicinity of the gauze and the duration of low-speed movement, which characterize the continuous state. The risk potential score is the result of the fusion of the time-varying risk base value and the continuous risk accumulation term.

[0114] Further explanation: In S5, the linkage control action that guides the excitation light source to adjust its excitation parameters specifically includes: periodically identifying gauze with an identity confidence level lower than a preset control threshold among all tracking instances as targets to be confirmed, and controlling the excitation light source to prioritize scanning the location area of ​​the targets to be confirmed.

[0115] In a preferred embodiment of the invention, a Python script checks the state vectors of all gauze instances maintained by ESKF at a fixed frequency of 10 times per second. The script reads the identity confidence threshold from a configuration file. For any instance with an identity confidence lower than the threshold, representing a target that has not been confirmed by fluorescence for an extended period, the script marks it as a "target awaiting confirmation." Subsequently, the script sends the corrected position coordinates of these targets to the controller of the multispectral excitation system, instructing it to concentrate the excitation energy and duration of the next round on these coordinate regions. When the system detects an interruption in the data stream from any target awaiting confirmation or a confidence level that remains below a lower degradation identity confidence threshold, it triggers a graceful degradation strategy, switches to pure visual tracking mode, and issues an alarm.

[0116] Further explanation: The linkage control actions also include:

[0117] Calculate a detection priority score for each target to be confirmed; and control the excitation light source to scan the area where the target is located in descending order of detection priority score.

[0118] Further explanation: The calculation of the detection priority score incorporates at least the following three technical elements:

[0119] The current confidence level of the target's identity; the rate of change of the confidence level over the past time window; and the risk potential score of the target.

[0120] The following is a detailed description of the implementation of the above content: Before the calculation process of this embodiment begins, the central control module will first load all configurable running parameters from the locally stored spreadsheet file;

[0121] The risk potential score, denoted as RPS, is a continuous floating-point number normalized to the [0,1] interval. It is used to quantify the potential legacy risk of a single piece of gauze at the current moment. 0 represents no risk, and 1 represents the highest risk level. RPS is calculated by the risk assessment module, and its value is the fusion result of the time-varying risk base value BRV and the continuous risk accumulation term CRA. The calculation model is as follows: calculate BRV and CRA, then add them together, and finally compare the sum with 1.0, taking the smaller one as the final RPS to ensure that its value does not exceed the upper limit.

[0122] The time-varying risk baseline value, denoted as BRV, is the event-driven component of RPS, used to respond to high-risk discrete events and instantly raise them to a higher baseline level. The specific determination method is as follows:

[0123] The event detection module is configured to monitor two Boolean conditions in parallel: Condition 1, gauze tracking is interrupted and its last position is within the "human suture area"; Condition 2, gauze enters the "instrument obstruction area"; if Condition 1 is true, the "interruption risk base value BRV" is obtained from the configuration unit. interrupt "As the output of BRV; if condition two is true, then obtain the "Occlusion Risk Base Value BRV" occlusion "As output; otherwise, the output is zero."

[0124] For a tracking trajectory to be classified as "interrupted," the necessary and sufficient condition is that the target's identity confidence level, within a sliding time window of 15 frames (corresponding to 0.5 seconds), is lower than the degradation threshold D. threshold This judgment method based on the average value of the time window can effectively filter out the brief drop in confidence caused by instantaneous occlusion or model fluctuations, and avoid erroneous "interruption" misjudgments.

[0125] Furthermore, to prevent frequent switching of risk states caused by the gauze lingering at the area boundary, this embodiment employs a judgment logic with hysteresis comparison. Specifically, an "entry" event is triggered when the gauze moves from outside the area to inside.

[0126] The "leave" event is triggered when the gauze moves to a safe margin beyond the boundary of the area; in this embodiment, the safe margin is set to 20 pixels; only after the "leave" event is confirmed can the next "enter" event be triggered again.

[0127] The determination of the sutured area, instrument-occluded area, and non-discarded area is as follows: In a specific worksheet of the spreadsheet file, each row represents a region. This row contains the region name and a series of vertex coordinate pairs (x, y) used to define the polygonal boundary of the region. During initialization, the spatial analysis module reads this worksheet, loads these vertex coordinates into memory, and constructs the geometric object used for subsequent spatial relationship determination. In a coordinate system with the top-left corner of the image as the origin (0,0) and the bottom-right corner as (1920, 1080), the "non-discarded area" is defined as a rectangular region composed of points (500, 200), (1400, 200), (1400, 800), and (500, 800).

[0128] The continuous risk accumulation term, denoted as CRA, is the cumulative effect component of RPS, used to simulate the slow accumulation and decay of risk over time. The risk assessment module determines the CRA through an iterative update process, its core idea derived from decay and accumulation models in physics. The calculation model for each time step is as follows:

[0129] Calculate the risk increment Risk at the current time step Increment The increment term is the regional proximity factor PZf plus the base cumulative coefficient C read from the configuration file. base The product of and . Obtain the continuous risk accumulator (CRA) from the previous time step. previous Then multiply it by a decay factor λ, which is less than 1; the decay factor λ is read from the configuration file. Add the calculated risk increment to the decayed historical cumulative term to obtain the continuous risk accumulation term CRA for the current time step.

[0130] The proximity factor, denoted as PZf, is a scalar in the range [0,1] that represents the proximity of the gauze to the nearest risk area that includes the "non-discard area". The closer the distance, the larger the PZf value.

[0131] The spatial analysis module is configured to: obtain the shortest Euclidean distance between the current position of the gauze and the boundary of the preset "non-discard area" polygon; and compare this shortest Euclidean distance with the "influence radius R" obtained from the configuration unit. effect The distance is compared with the radius of influence; if the distance is greater than the radius of influence, then PZf is zero; otherwise, the PZf value between 0 and 1 is obtained by subtracting the ratio of the distance to the radius of influence from 1.

[0132] It should be noted that: BRV interrupt BRV occlusion C base ,λ,R effect The determination method is as follows: for BRV interrupt BRV occlusionA validation dataset, ethically reviewed and anonymized, was prepared. This dataset contains 250 independent video clips, all in color at 1920×1080 resolution and 30fps, with an average duration of 15 seconds. 100 videos explicitly demonstrate "gauze tracking loss within the suture area," 100 explicitly demonstrate "gauze completely covered by surgical instruments," and the remaining 50 serve as a control group with normal procedures. The dataset covers at least three different surgical procedures and includes various lighting conditions, including normal illumination, strong surgical field light, and shadow. An expert panel of at least 10 senior surgeons with over 10 years of surgical experience was invited. The video clips were randomly presented to each expert, who was asked to immediately rate the "clinical risk level" of the event after viewing, ranging from 0 (no risk) to 10 (extremely dangerous). All expert ratings were collected. The mean scores for the "tracking loss" event group and the "instrument coverage" event group were calculated separately. These two mean scores were then linearly transformed to the [0,1] interval. The average score of the mapped "tracking loss" event group is then determined as the interruption risk baseline value (BRV). interrupt The final calibration value; the average score of the "device coverage" event group is determined as the occlusion risk baseline (BRV). occlusion The final calibration value.

[0133] Regarding the basic cumulative coefficient C base Attenuation factor λ, radius of influence R effect Determination:

[0134] Design an objective function to quantify the difference between the Risk Potential Score (RPS) curve and the expert expectation curve. Specifically, prepare a set of typical risk evolution scenarios:

[0135] Scenario 1 (Slow Approach and Stop): Simulates gauze moving at a speed of 5 pixels / second from an area of ​​influence radius R. effect It moves at a constant speed in a straight line towards the center of the "non-discard area" and remains stationary at the center point for 30 seconds.

[0136] Scene 2 (Fast Crossing): Simulates gauze moving horizontally across the "non-discard area" at a speed of 50 pixels per second without stopping.

[0137] Scenario 3 (Wandering at the Edge): Simulates the gauze moving randomly within a small range inside and outside the boundary of the "non-discard area".

[0138] The expert panel was then asked to plot the curves showing the ideal RPS over time, which corresponded to their expected performance. The objective function used was root mean square error, which was employed to calculate the area difference between the RPS curve output by the model and the curve expected by the experts.

[0139] The Particle Swarm Optimization (PSO) algorithm is used for parameter optimization. The basic cumulative coefficient C is... base Attenuation factor λ, radius of influence R effect As a three-dimensional particle to be optimized, in each iteration, a risk assessment model is run using a set of parameters to simulate the typical scenario described above, and the error between the output RPS curve and the expert expectation curve is calculated. After at least 10 iterations, the set of parameters (C) that minimizes the objective function globally is selected. base ,λ,R effect The parameter combination is used as the optimal parameter value for final solidification in this embodiment.

[0140] In this embodiment, the risk threshold BRV is used. interrupt The calibration value is 0.85; the baseline value for occlusion risk (BRV) is [value missing]. occlusion The calibration value is 0.70. These two values ​​reflect the expert panel's assessment that "tracking interruption in the suture area" is a higher-risk event than "instrument obstruction."

[0141] The dynamic parameter was determined by optimizing the particle swarm optimization algorithm as: the basic cumulative coefficient C. base The preferred value is 0.05; the preferred value for the attenuation factor λ is 0.98, which ensures that the risk accumulation has a certain memory effect without growing indefinitely; the radius of influence R effect The preferred value is 150 pixels, which corresponds to a physical distance of approximately 10 centimeters at 1080p resolution.

[0142] The detection priority score, denoted as PPS, is a dimensionless floating-point number used to rank all targets to be confirmed. A higher score indicates a higher priority for scanning. The active detection scheduling module calculates the PPS for each target to be confirmed. Its calculation model is a weighted sum, inspired by the multi-attribute utility function in decision theory: parallel calculation of three components normalized to the [0,1] interval.

[0143] Inverse confidence transformation component CIc: Subtract the current identity confidence from 1.0;

[0144] Confidence change rate component CRc: Calculates the linear regression slope of identity confidence within the past time window of "1 second". If it is negative, take its absolute value and then normalize it.

[0145] Risk-related component RRc: The RPS value calculated by the risk assessment module is directly used.

[0146] Read three weighting coefficients from the configuration file: confidence weight wc, rate of change weight wr, and risk weight w. riskThe final value of the detection priority score (PPS) is obtained by multiplying CIc by wc, CRc by wr, and RRc by wc. risk Multiply them, and then add the three products together.

[0147] It should be noted that: (wc,wr,w) risk This determines the trade-offs among the three decision dimensions: "uncertainty," "the worsening trend of uncertainty," and "potential risk." The optimal value must be determined through system-level, performance-oriented optimization experiments; the specific method is as follows: find a set of optimal weight coefficients (wc, wr, w...). risk This minimizes the average response time of the system when solving "high-risk, high-uncertainty" problems on the validation dataset.

[0148] Risk-weighted uncertainty elimination time (RWURT) is defined. For each "target to be confirmed" in the test set, this metric is calculated as follows: the time taken from when the target is first marked as "target to be confirmed" until its identity confidence recovers to above the threshold is multiplied by the target's average RPS value during this period. The final score for the entire test set is the average of the RWURTs for all such events. The smaller the metric, the better the performance. A grid search combined with cross-validation is employed. For each combination, the full-process method of this invention is run on a large labeled video validation dataset containing various complex scenarios, and its final RWURT score is calculated.

[0149] Choose the weight combination (wc,wr,w) that minimizes the global RWURT score. risk The optimal parameter values ​​for final solidification in this embodiment are: confidence weight wc, with a preferred value of 2; rate of change weight wr, with a preferred value of 3; and risk weight wc. risk Its preferred value is 5.

[0150] If a risk alerting system is based on isolated, hard-coded rules, it cannot handle the ambiguity and dynamic evolution of risks, easily leading to missed or false alarms. Meanwhile, conventional scanning systems employ polling or random strategies, lacking an understanding of the semantics of the scenario, resulting in wasted resources and response delays. This invention constructs a complete information processing and control closed loop, from "state awareness" to "risk quantification" and then to "priority scheduling."

[0151] Step S4, characterized by a risk quantification layer, transforms risk from a Boolean value into a continuous physical quantity with "memory" and "attenuation" characteristics by fusing the event-driven time-varying risk baseline (BRV) and the process-driven continuous risk accumulation term (CRA). This enables the differentiation between "instantaneous high risk" and "continuously accumulating potential risk," providing richer and more refined input for subsequent decision-making.

[0152] Step S5 characterizes the priority scheduling layer S5: a weighted fusion model of PPS, unifying three different dimensions of information—the confidence inverse transformation component CIc representing the severity of the current state, the confidence rate of change component CRc representing the speed of state deterioration, and the risk correlation component RRc representing the severity of the state's consequences—within the decision framework. This allows for prioritizing high-risk but stable-confidence targets over low-risk but slightly weaker-confidence targets. This mechanism possesses the capability of "optimal resource allocation under risk perception," ensuring that the scanning time of the excitation source is used to address the most significant and deterministic threat to the overall system security.

[0153] In this embodiment, the complete calculation process of risk assessment and linkage control is scheduled and executed by the central control module, specifically including the following steps: The risk assessment module and the active detection scheduling module continuously receive corrected state vectors containing absolute identity, corrected motion trajectory, and identity confidence for all gauze instances from the upstream entangled state Kalman filter model (ESKF). It should be noted that in the corrected state vector, absolute identity is ultimately represented as a unique identifier (UID) matched and assigned from the database. The "identity confidence" quantifies the degree of certainty of the ESKF regarding the matching relationship between the "currently observed fluorescent fingerprint" and the "assigned UID".

[0154] For each gauze instance, the risk assessment module is executed in parallel:

[0155] Calculate its regional proximity factor PZf. Update its continuous risk accumulation term CRA. The event detection model determines its time-varying risk baseline BRV. The fusion model calculates the final risk potential score RPS and updates the state vector of this instance with this score.

[0156] The active detection scheduling module initiates a scheduling cycle at a preset frequency of 10 times per second. It filters out all identities with a confidence level below the identity confidence threshold C. threshold Examples are used to form a "targets to be confirmed" list. In this embodiment, the identity confidence threshold is represented as a preset control threshold; the identity confidence threshold C threshold The preferred example value is 0.7. Choosing this value involves a trade-off: a value that is too high will cause the system to initiate active probing too frequently, wasting resources; a value that is too low may lead to delayed responses to truly uncertain targets. 0.7 is the balance point that ensures system sensitivity while avoiding excessive resource consumption.

[0157] For each target in the list, its detection priority score (PPS) is calculated. The "targets to be confirmed" list is sorted in descending order of PPS. The corrected position coordinates of the targets in the sorted list are extracted, and a scanning command with priority information is generated. The scanning command is sent to the controller of the multispectral excitation system, which executes the prioritized, focused scanning action. To ensure better robustness, this embodiment also includes an emergency response and degradation mechanism. This mechanism is continuously monitored and executed by the active detection scheduling module, and its specific implementation is as follows:

[0158] Triggering condition: The active detection scheduling module maintains a detection count for each "target to be confirmed". When the confidence level of a target's identity continuously falls below the preset degradation threshold D... threshold Furthermore, the cumulative number of times it has been prioritized for detection exceeds the maximum number of detections N. max At that time, the downgrade mechanism is triggered.

[0159] Degradation threshold D threshold Its physical meaning is the minimum confidence level limit that can be tolerated. In this embodiment, its preferred value is 0.4. This value was determined experimentally; a value lower than this indicates that the target state is unreliable.

[0160] Maximum number of detections N max Its physical meaning is the maximum resources that the target is willing to expend. In this embodiment, its preferred value is 5 times.

[0161] When the degradation mechanism is triggered, the following actions will be taken for the target: the tracking mode of the target will be forcibly switched from "fusion tracking" to "pure visual tracking," and fluorescence excitation will be stopped to conserve excitation light source resources. In the target's state vector, its Risk Potential Score (RPS) will be forcibly set to the highest value of 1.0. Through the user interface, a high-priority visual alarm containing the target's last location coordinates will be generated, and its status will be explicitly marked as "signal lost, manual confirmation required," until the operator manually clears the alarm. Prioritizing the scanning of the excitation light source increases the probability of acquiring the fluorescent fingerprint of low-confidence targets. Once successfully acquired, ESKF will update its identity confidence with high weight. The increased confidence will prevent it from being listed as a "target awaiting confirmation" in the next round of scheduling, and its associated PPS will also decrease, thus forming a complete intelligent closed loop of "problem discovery → problem assessment → priority resolution → problem elimination."

[0162] The following is a detailed implementation description of the above content: This technical solution generates two core quantitative outputs: Risk Potential Score (RPS) and Detection Priority Score (PPS).

[0163] The higher the risk potential score, the higher the potential residual risk of the gauze instance and the more serious its deviation from the safety baseline. Specifically, when the risk potential score (RPS) is closer to 1, it indicates that the potential residual risk of the gauze instance is higher and its state deviates more seriously from the safety baseline. Conversely, when the RPS is closer to 0, it indicates that the movement trajectory of the instance is more in line with safety expectations and the residual risk is lower.

[0164] A higher detection priority score (PPS) indicates that the target to be identified needs to be prioritized; the urgency of guiding the excitation source to scan it and the degree of resource allocation bias are also stronger. Conversely, a lower PPS indicates that the identification requirement of the target is relatively less important and can be placed later in the resource scheduling sequence.

[0165] Analysis of the impact of Risk Potential Score (RPS): Time-varying Risk Base Value (BRV) and Continuous Risk Accumulation Term (CRA): Both are positively correlated with the final RPS. RPS is designed as a fusion of these two risk components, aiming to comprehensively capture the overall risk constituted by both unexpected events and ongoing conditions.

[0166] Regional proximity factor PZf: When other parameters remain constant, an increase in PZf, meaning the gauze is closer to the risk area, will lead to a monotonically increasing RPS by increasing the risk increment of the continuous risk accumulation term CRA. This aligns with the logic that "the closer to the hazard source, the higher the potential risk."

[0167] The attenuation factor λ (0 < λ < 1) is positively correlated with the historical RPS value. The larger λ is, the stronger the CRA's "memory" of historical risk, and the slower the risk attenuation. This allows for adaptation based on the risk characteristics of different surgeries; fast-paced surgeries require rapid risk attenuation, while slow-paced surgeries require the opposite.

[0168] Analysis of the impact of the detection priority score (PPS): It shows a negative correlation with PPS through the "inverse confidence transformation component (CIc)". That is, the lower the identity confidence, the larger the CIc, and the higher the final PPS.

[0169] Identity confidence change rate: It is negatively correlated with PPS through the "confidence change rate component CRc". That is, the smaller the confidence change rate, the larger the CRc, and the higher the final PPS.

[0170] Risk Potential Score (RPS): The risk-related component RRc is positively correlated with PPS. This ensures that attention is paid not only to "uncertainty" but also to "uncertainty with risk." The confirmation requirement for high-risk targets is far higher than that for low-risk targets, even if both have the same confidence level. This allows limited detection resources to be precisely directed to the intersection of risk and uncertainty.

[0171] To quantify and verify the beneficial effects of the technical solution of this invention, the following representative surgical scenarios were designed, and the performance of the "conventional method" and the "method of this invention" on key indicators were compared: This experiment aims to verify the significant progress of this invention in the accuracy of risk assessment and the efficiency of active detection. S4 uses a simple binary risk label (0 or 1), while the conventional method S5 uses polling scanning (without priority) for all low-confidence targets. This method, however, uses quantified RPS and priority-driven PPS. By comparing the RPS, PPS, and the final risk-weighted uncertainty elimination time (RWURT) under four typical scenarios, the core advantages of this invention can be clearly revealed; see below for details. Figure 1 As shown:

[0172] Table 1: Performance Comparison Experiments of Risk Assessment and Detection Scheduling in Different Scenarios

[0173] Parameter name Scenario 1: Low confidence in the safe zone Scenario 2: Hovering on the edge of the risk zone Scenario 3: Interruption of the suture track Scenario 4: Confidence level in the risk zone drops sharply Scenario 5: Confidence level in the risk zone gradually decreases Scenario 6: Dual-Objective (A5 / B5) Competition Identity confidence 0.6 0.9 1.0 (before the interruption) 0.8→0.5 0.8→0.75 0.6 / 0.6 Confidence level change rate ( / second) 0 0 N / A -0.3 -0.05 0.0 / 0.0 Event triggered none none Interruption of suture area none none None / None Regional proximity factor 0.1 0.8 N / A 0.9 0.9 0.9 / 0.2 Risk indicator (0 / 1) 0 0 1 0 0 0 / 0 Detection priority middle N / A N / A middle middle Same (polling) Risk potential score 0.05 0.64 0.9 0.72 0.72 0.72 / 0.08 Detection priority score 1.05 0.72 N / A 6.4 4.4 4.4 / 1.6 RWURT (seconds) - Standard Method 5 N / A N / A 12 12 12.0 (Target A5) RWURT (seconds) - This invention 2.5 N / A N / A 3 6.5 4.0 (Target A5)

[0174] Risk-weighted Uncertainty Elimination Time (RWURT): This is the core performance indicator set in this experiment, used to comprehensively evaluate the efficiency of solving key problems. It is calculated as follows: the time from when a target is identified as "pending confirmation" until its confidence level returns to normal, multiplied by the average risk potential score (RPS) of that target during this period. The lower the value, the faster the system can resolve more dangerous problems.

[0175] Comparing Scenario 2 with conventional methods: In Scenario 2, "hovering on the edge of the risk zone," the gauze did not trigger any binary risk rules in conventional methods, therefore its "risk flag" was 0, and the system could not perceive the potential risk. However, the RPS of this invention, through the mechanism of continuous risk accumulation terms, accurately quantifies a potential risk value as high as 0.64. This proves that compared with conventional rule judgments based on discrete events, this invention can identify potential risks accumulated from continuous states earlier and with greater precision, improving the sensitivity and foresight of risk perception.

[0176] Comparing Scenario 1 and Scenario 4: In Scenario 1, the target confidence level is low (0.6) but the risk-associated component RRc is low (RPS=0.05). In Scenario 4, the target confidence level is relatively high (minimum 0.5) but the risk-associated component RRc is high (RPS=0.72) and is rapidly deteriorating.

[0177] Conventional methods, unable to perceive risk, assign similar detection priorities to both scenarios. However, the PPS calculation results of this invention show that the PPS for scenario four (6.4) is higher than that for scenario one (1.05). This demonstrates the effectiveness of the detection priority score (PPS) design of this invention, which successfully and decisively allocates detection resources from a "low-risk, high-uncertainty" target (scenario one) to a "high-risk, high-deterioration-trend" target (scenario four).

[0178] Comparing the RWURT in Scenario 4: In the most critical Scenario 4, "sudden drop in confidence in the risk area," the conventional method, due to its indiscriminate polling scan, has an RWURT of 12.0 seconds. This invention, through the PPS mechanism, immediately places the target in the highest scan priority, reducing its RWURT to 3.0 seconds. The performance improvement is calculated as: [(12.0-3.0) / 12.0]×100%=75%.

[0179] In the control group consisting of Scenario 5 and Scenario 4, the target's risk potential score (RPS) was exactly the same (both 0.72), but the confidence level in Scenario 4 "dropped sharply" (rate of change -0.3), while in Scenario 5 it "dropped slowly" (rate of change -0.05). Data shows that the PPS of Scenario 4 (6.4) was higher than that of Scenario 5 (4.4).

[0180] This comparison demonstrates that, even under the same risk level, the PPS calculation mechanism of this invention can still assign higher processing priority to objects whose state deteriorates faster by sensing the dynamic trend of "confidence change rate." This directly and strongly supports the technical feature regarding the integration of the "change rate" element.

[0181] Scenario 6 simulates a common multi-objective competition situation in clinical practice. Target A5 and target B5 have the same low confidence level, but target A5 is in the high-risk region (RPS=0.72), while target B5 is in the safe region (RPS=0.08).

[0182] Conventional methods, unable to differentiate risks, perform indiscriminate polling scans on both targets. This invention, however, accurately calculates that target A5's PPS (4.4) is significantly higher than target B5's PPS (1.6), thus prioritizing the processing of the truly critical target A5. Its RWURT (Recovery Timer) is 4.0 seconds, a 66.7% improvement over the conventional method's 12.0 seconds. The calculation is: [(12.0-4.0) / 12.0]×100%, further demonstrating the robustness and efficiency of this invention in complex scenarios.

[0183] To translate the quantitative output of this invention into explicit, actionable system behaviors and clinical procedures, the following application ranges are defined. Figure 2 and Figure 3 :

[0184] Figure 2 Application range of Risk Potential Score (RPS)

[0185] Risk level RPS range Behavioral / Clinical Recommendations Safety [0,0.3] The gauze instance silently records trajectory data in the background and is displayed in the user interface in the standard color (green). Warning (0.3,0.7] The gauze instance is highlighted (yellow) in the user interface and accompanied by a non-invasive audio prompt. Nurses are advised to monitor the current status of the gauze. Danger (0.7,1.0] The gauze instance flashes continuously in a bright red color in the user interface, triggering a high-level, non-ignorable audible alarm. This is logged as a high-risk event in the system log.

[0186] Figure 3 Application range division of detection priority score (PPS)

[0187] Priority PPS range (example) Excitation Light Source Linkage Control Strategy conventional [0,2.0] Place the target at the end of the standard scan queue and perform polling scans at the normal frequency. priority (2.0,6.0] The target is placed in a priority queue, the scanning frequency is increased to twice the normal rate, and the single excitation energy is appropriately increased. urgent (6.0,10] Immediately interrupt all current routine scanning tasks and focus all detection resources (highest frequency, maximum energy) on the area where the target is located until its confidence level is restored or it is manually confirmed.

[0188] The computational logic involved in this application can be constructed using algorithms such as regression analysis in machine learning, establishing a mathematical model by analyzing the inherent trends and interrelationships of the collected parameters. This process can be implemented using specialized computational tools (such as Python's Scikit-learn library or the R language environment). Throughout all calculations, to eliminate the influence of different physical dimensions and ensure that data is compared and analyzed on the same scale, the input parameters in each formula are dimensionless. The dimensionless techniques used include, but are not limited to, max-min normalization or Z-score standardization.

[0189] The algorithm of this invention is implemented as a Python script. Before executing the core logic, the program first executes a data loading module (e.g., using the widely used pandas library in Python) configured to read the aforementioned spreadsheet file and load its contents into the program's working memory (e.g., a DataFrame data structure). Subsequent algorithm steps will directly query and retrieve the required configuration parameters from this in-memory data structure.

[0190] It should be emphasized that the foregoing embodiments are merely illustrative of preferred implementations of the present invention and are not intended to limit the scope of protection of the present invention. This application also provides a computer-readable storage medium having computer program instructions stored thereon.

Claims

1. A method for in-surgery gauze counting and tracking based on visual perception, characterized in that, The specific steps include: S1: Obtain a multi-spectral image sequence, and extract first data for representing the inherent physical identity of at least one piece of gauze from the multi-spectral image sequence, wherein the first data is a fluorescence fingerprint vector; The fluorescence fingerprint vector is obtained by: sequentially exciting the gauze through a multi-spectral excitation system, and synchronously collecting fluorescence response images generated under different excitation wavelengths, thereby constructing the fluorescence fingerprint vector representing the response characteristics of the gauze in a multi-dimensional spectral space; S2: Obtain a visual image sequence, and process the visual image sequence through a space-time tracking model to obtain second data for representing the dynamic motion trajectory of at least one piece of gauze, wherein the second data includes a motion state representing the position and speed of the gauze; S3: Based on a preset entangled state Kalman filter model, the first data and the second data are fused and processed, wherein the fusion processing includes: jointly estimating the motion state of the gauze and the identity confidence corresponding to the fluorescence fingerprint vector in a unified state vector, and periodically correcting the identity confidence using the first data to suppress the error accumulation of the space-time tracking model; S4: Based on the result of the fusion processing, a corrected state vector representing the absolute identity of the gauze and the corrected motion trajectory is generated; S5: Based on the corrected state vector, a linkage control action is performed, which at least includes: feeding back the corrected motion trajectory to the space-time tracking model to calibrate its tracking reference, and when the identity confidence in the corrected state vector is lower than a preset control threshold, guiding the excitation light source to adjust its excitation parameters for the gauze.

2. The method of claim 1, wherein: In S2, the space-time tracking model is a space-time graph neural network model, which is configured to: take the gauze candidate targets in different video frames as graph nodes, construct the space-time correlation edges between the graph nodes, and infer and maintain the identity continuity and motion trajectory of each piece of gauze in the video sequence by propagating and aggregating node information on the graph structure.

3. The method of claim 2, wherein: In S3, the entangled state Kalman filter model performs the following in its update step: When the first data is obtained, the matching degree of the fluorescence fingerprint vector with each record in the preset fingerprint database is calculated, and the identity confidence component in the state vector is updated with a first weight based on the matching degree; When the first data cannot be obtained, the identity confidence component is updated with a second weight smaller than the first weight based on the visual appearance features of the gauze; The greater the value of the identity confidence component, the higher the certainty level of the identity estimation of the entangled state Kalman filter model for the current gauze.

4. The method of claim 3, wherein: Before the update step, it also includes: determining a fluorescence signal quality factor representing the signal quality of the first data; and, determining a visual tracking stability factor representing the tracking stability of the second data; In the update step, the first weight and the second weight are dynamically calculated based on the fluorescence signal quality factor and the visual tracking stability factor; The first weight value is positively correlated with the value of the fluorescence signal quality factor, and the second weight value is positively correlated with the value of the visual tracking stability factor.

5. The method of claim 4, wherein: In S4, the corrected state vector further comprises a risk indicator for representing a risk level of the gauze left behind; The risk indicator is determined based on analysis of the corrected motion trajectory of the gauze, and the analysis comprises: detecting whether the gauze is continuously stationary in a preset non-discarding area for more than a first time threshold, or whether the tracking trajectory of the gauze is interrupted in a human body suture area.

6. The method of claim 5, wherein: The risk indicator is a risk potential score; The determination of the risk potential score comprises: determining a time-varying risk base value based on discrete events of whether the gauze enters a preset risk area or whether the tracking trajectory of the gauze is interrupted; and determining a continuous risk accumulation item based on a continuous state represented by the degree of the gauze adjacent to the risk area and the duration of low-speed motion; The risk potential score is a fusion result of the time-varying risk base value and the continuous risk accumulation item; The greater the risk potential score, the higher the potential risk of the gauze left behind is represented, and the more serious the state of the gauze deviates from the safety baseline.

7. The method of claim 6, wherein: In S5, the linkage control action of guiding the excitation light source to adjust its excitation parameters comprises: periodically identifying gauzes with an identity confidence lower than a preset control threshold in all tracking instances as to-be-confirmed targets, and controlling the excitation light source to preferentially scan a position area where the to-be-confirmed targets are located.

8. The method of claim 7, wherein: The linkage control action further comprises: calculating a detection priority score for each to-be-confirmed target; the greater the value of the detection priority score, the more the to-be-confirmed target needs to be preferentially processed is represented; and controlling the excitation light source to scan the position area where the to-be-confirmed targets are located in the order from high to low of the detection priority scores.

9. The method of claim 8, wherein: The calculation of the detection priority score fuses at least the following three technical elements: the current identity confidence of the to-be-confirmed target; the change rate of the identity confidence in a past time window; and the risk potential score of the to-be-confirmed target.

Citation Information

Patent Citations

  • Surgical foreign matter leaving monitoring method and device, computer equipment and storage medium

    CN120747867A

  • Medical clinical operation skill auxiliary evaluation method and system based on artificial intelligence

    CN118628947A

  • MRI-based augmented reality assisted real-time surgery simulation and navigation

    US20230114385A1