Wearable smart glasses system and control method thereof
Patent Information
- Application Number
- CN202611039106.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]有鉴于此,本发明提供了一种佩戴式智能眼镜系统及其控制方法,解决现有神经外科手术的耗材选型配合方式,存在沟通效率低、精准度差、手术风险高、无菌安全性不足等诸多缺陷,难以满足精细化、高效化、安全化的神经外科手术临床需求
[0016]经由上述的技术方案可知,与现有技术相比,本发明公开提供了一种佩戴式智能眼镜系统及其控制方法,具备架构轻量化、算法鲁棒性强、临床适配度高、术中安全性优的核心技术优势,有效克服了现有神经外科手术中医护口头沟通滞后、耗材选型主观化、术区无菌风险高、创面尺寸量化缺失的问题。
Smart Images

Figure CN122827802A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart glasses technology, and more specifically to a wearable smart glasses system and its control method. Background Technology
[0002] Neurosurgery is a highly technical medical procedure primarily used to treat various neurological diseases, such as brain tumors, cerebrovascular diseases, and traumatic brain injury. With the continuous advancement of medical technology, neurosurgical techniques are constantly evolving and improving. Microscopy is one of the most important tools in neurosurgery. Through a surgical microscope, surgeons can magnify the surgical field, observe the nervous system structures more accurately, and thus remove diseased tissue more precisely.
[0003] In neurosurgical hemostasis procedures, encephalosporin pads and gelatin sponges are the main hemostatic consumables, and the precise selection of their specifications and dimensions directly determines the hemostatic effect and the efficiency of surgical progress. However, in current clinical surgical scenarios, scrub nurses cannot directly view the microscope imaging field and cannot intuitively obtain key information such as the actual wound size and bleeding range in the surgical area. They can only rely on their own clinical experience and verbal communication with the surgeon to subjectively judge the required specifications and dimensions of encephalosporin pads and gelatin sponges, making it difficult to guarantee the accuracy and timeliness of consumable selection.
[0004] Based on microscopic observation of the surgical area, the surgeon verbally announces the specifications and models of the necessary consumables, and the scrub nurse retrieves the corresponding supplies. This communication method has significant drawbacks. Intraoperative verbal announcements carry multiple risks, including information delays, misunderstandings by medical staff, and mishandling or incorrect use of consumables. This not only directly reduces the efficiency of intraoperative hemostasis and delays the optimal time for hemostasis but also significantly increases surgical risks, impacting overall surgical safety. Furthermore, the use of vague terms like "small," "medium," and "large" to define consumable specifications during clinical practice lacks standardized, objective quantitative criteria. Individual differences in surgeons' understanding and judgment of consumable sizes lead to discrepancies between the specifications announced by the surgeon and the nurse's perception, further exacerbating consumable selection errors and severely hindering the standardization of surgical procedures.
[0005] To avoid selection errors, scrub nurses often need to frequently look down at the instrument table and repeatedly compare the actual dimensions of various cotton pads and gelatin sponges to confirm the specifications of consumables. This practice not only significantly increases the time spent on unnecessary pauses during surgery, prolonging the overall surgical procedure and reducing surgical efficiency, but more importantly, the nurse's repeated bending down and raising of the hand to compare increases the probability of contact between sterile and non-sterile areas, easily causing contamination of the surgical area and increasing the patient's risk of postoperative infection.
[0006] Therefore, in view of the shortcomings of existing technologies, how to provide a wearable smart glasses system and its control method to meet the clinical needs of precise, efficient and safe neurosurgical operations is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a wearable smart glasses system and its control method, which solves the problems of existing neurosurgical consumable selection and matching methods, which have many defects such as low communication efficiency, poor accuracy, high surgical risk, and insufficient aseptic safety, making it difficult to meet the clinical needs of refined, efficient and safe neurosurgical operations.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: a wearable smart glasses system, comprising: a smart glasses body, an image acquisition module, an image preprocessing module, a reference object recognition module, a size calculation module, and a display module; The smart glasses integrate a camera and a near-eye display optical module for capturing microscope images; The image acquisition module is used to acquire microscope images in real time; The image preprocessing module is used to preprocess the microscope image; The reference object recognition module has a built-in trained medical image recognition model, which is used to recognize various standard reference objects of known fixed size in the surgical field based on preprocessed microscope images. The size calculation module is used to establish a mapping relationship between the image pixel size and the actual physical size of the standard reference object, calculate the actual area and boundary size of the bleeding area and tissue defect area in the surgical area, and automatically convert the appropriate surgical consumables specifications by combining the matching algorithm. The display module is used to project the calculated specifications of surgical consumables onto the eyeglass lenses for display in real time.
[0009] Preferably, the image preprocessing module includes a pre-denoising unit, a perspective distortion correction unit, an RGB-HSI conversion unit, a hybrid threshold segmentation unit, a grayscale conversion unit, a multi-resolution feature decomposition unit, and a feature vectorization unit connected in sequence. The multi-resolution feature decomposition unit is used to decompose a grayscale image into a high-resolution feature map HR and a low-resolution feature map LR. The low-resolution feature map (LR) is scaled using bilinear interpolation to obtain a size-matched medium-resolution feature map (MR), and gradient texture detail features are extracted from the medium-resolution feature map (MR). ; Based on the grayscale and luminance components of the high-resolution feature map HR Gray-scale luminance components of the medium-resolution feature map (MR) High-frequency detail features of the image were extracted. ; The feature vectorization unit is used to analyze high-frequency detail features of the image. Gradient texture detail features The feature blocks are divided into blocks and converted into high-dimensional image vectors X and low-dimensional image vectors Y, which are then used as the output standardized feature vectors.
[0010] Preferably, the hybrid threshold segmentation unit employs a saturation-based dual-threshold weighted fusion algorithm, as shown in the following formula: ; in, This represents the final adaptive saturation threshold used for binary segmentation of the wound. This represents the adaptive segmentation threshold, which is solved by maximizing the inter-class variance between foreground and background. The minimum error segmentation threshold is defined by minimizing the pixel classification misclassification error. This represents the fixed-weight hyperparameter of the weighted fusion; The binary segmentation determination rule is based on pixel saturation. Pixels identified as foreground pixels of bleeding / wound are marked as mask ground truth value 1; Pixels identified as surgical instruments or normal tissue backgrounds are marked as mask ground truth value 0, and a binary mask image is output.
[0011] Preferably, the medical image recognition model adopts an improved NAM-YOLACT algorithm, including: an input layer, a lightweight residual backbone layer, a NAM normalized attention enhancement layer, a multi-scale FPN fusion layer, a dual-branch inference layer, and a processing output layer; The input layer receives the preprocessed microscope image; The lightweight residual backbone layer adopts the ResNet-50-DW bottleneck backbone network and sets four residual stages to extract image features step by step, respectively extracting surface texture features, mid-layer geometric contour features, and deep semantic features of the surgical field. The NAM normalized attention enhancement layer is embedded in the output of each level of the lightweight residual backbone layer and the back end of the multi-scale FPN fusion layer, and includes two parallel branches: channel attention and spatial attention. The multi-scale FPN fusion layer constructs a five-level feature pyramid structure, and the fusion process embeds high-frequency detail features and gradient texture features output by the image preprocessing module. The dual-branch inference layer includes two parallel sub-branches: a detection and regression branch and a prototype mask generation branch. The detection and regression branch is used to output the target class probability, bounding box offset, and global matching confidence. The prototype mask generation branch is used to generate medical prototype mask basis vectors and generate instance masks through feature weighted combination. The processing output layer eliminates mismatched feature pairs by normalizing the correlation coefficient using LACC, and finally outputs the reference object category, sub-pixel level bounding box coordinates, and the total number of effective coverage pixels M of the reference object in a structured output.
[0012] Preferably, in the NAM-normalized attention enhancement layer, the channel attention output is: ; Spatial attention output: ; Channel / spatial weight normalization allocation formula: ; in, This represents the channel attention weight matrix; Represents the spatial attention weight matrix; Represents a sigmoid nonlinear activation function; This represents the batch normalization operator; The network intermediate layer feature maps representing the input to the attention module correspond to the channel branch and spatial branch inputs, respectively. This represents the global pixel mean of all feature maps within a single training batch; This represents the global variance of pixels across all feature maps within a single training batch. This represents the normalized learnable scaling factor; Represents the normalized learnable bias coefficients; Represents a numerically stable minimum constant; This represents the normalized channel and spatial attention weight coefficients; Represents the original learned weights for the i-th channel / spatial location; This represents the total number of channels and the total number of spatial locations in the feature map; This represents the Hadamard product of a matrix.
[0013] Preferably, the size calculation module establishes a pixel-to-real physical size mapping based on the identified standard reference object, uses eight-neighborhood tracking of the wound contour to solve for the bleeding / defect area, and graded matching of surgical consumables.
[0014] Preferably, the length and width of the standard reference object are compared with the microscope's real-time magnification factor k based on its inherent factory reference dimensions. , Magnification correction is performed to obtain the equivalent true length and width of the reference object adapted to the current microscope field of view. , ; Substituting the total number of pixels M effectively covered by the reference object into the pixel equivalent formula Calculate the actual physical area corresponding to a single pixel. ; A dynamic pixel-to-millimeter mapping relationship is established, and parameters are updated in real time when the microscope magnification is switched during surgery. After the mapping benchmark is completed, the complete edge of the wound is extracted using an eight-neighbor counterclockwise contour tracking algorithm based on the binary mask of the bleeding wound output by the image preprocessing module. The invalid pseudo-boundaries formed by breakpoints, noise, and occlusion are filtered out by an eight-direction vector encoding judgment rule, retaining only the valid contour points. The pure valid pixels inside the wound are counted. All pixels of the outline boundary ; Using the half-area correction formula Calculate the equivalent total pixels, and then combine them with the already solved actual physical area. The measured area of the wound was obtained: ; At the same time, the actual physical area The square root yields the actual side length coefficient of a single pixel. The actual maximum length and width of the wound are calculated by combining the extreme pixel coordinates of the wound outline. Based on the measured area and dimensions of the wound, the system automatically matches the corresponding specifications of gelatin sponge, brain cotton pads, and estimated usage quantities by retrieving locally solidified neurosurgical clinical grading thresholds. A fixed safety redundancy coefficient ξ is then introduced, and a proportional conversion formula is used. The system calculates the cutting dimensions of the consumables and finally sends the wound length and width, actual area, consumable model, estimated usage, and consumable cutting reference dimensions to the display module.
[0015] Preferably, a control method for wearable smart glasses includes: real-time acquisition of microscope images; The microscope images are preprocessed; Identify various standard reference objects of known fixed size within the surgical field based on preprocessed microscope images; A mapping relationship between image pixel size and real physical size is established for the standard reference object. The actual area and boundary size of the bleeding area and tissue defect area in the surgical area are calculated. The appropriate surgical consumable specifications are automatically converted by the matching algorithm. The calculated specifications of surgical consumables are projected onto the glasses lenses for display in real time.
[0016] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a wearable smart glasses system and its control method, which has the core technical advantages of lightweight architecture, strong algorithm robustness, high clinical adaptability and excellent intraoperative safety. It effectively overcomes the problems of lagging verbal communication between doctors and nurses in existing neurosurgical operations, subjective selection of consumables, high risk of sterility in the surgical area, and lack of quantitative wound size.
[0017] This invention constructs an embedded image processing architecture, which, through multi-level image preprocessing, an improved NAM-YOLACT medical instance segmentation network and a dynamic pixel-physical size mapping mechanism, relies on the real-time magnification of the microscope to dynamically correct the pixel equivalent, completely eliminating the system error of size measurement under zoom conditions. Combined with eight-neighbor contour tracking and boundary half-area correction strategies, it solves the technical shortcomings of traditional naked-eye assessment and subjective prediction without quantitative standards, and provides objective and accurate data support for consumable selection. This invention abandons the traditional intraoperative verbal communication mode and uses AR near-eye visualization display technology to push wound parameters, consumable specifications, dosage and cutting reference lines to the instrument nurse's field of vision in real time. This enables seamless synchronous transmission of surgical field information and consumable instructions, eliminating the risks of information transmission delay, mishearing human voice, and mistaking or missing consumables, and significantly shortening the time spent on hemostasis. The improved NAM-YOLACT model achieves real-time inference on edge embedded devices while maintaining millimeter-level fine-grained recognition accuracy through normalized attention enhancement and parameter sparsity compression. It does not rely on cloud computing power, eliminates network latency and disconnection risks, and is suitable for long-term, high-intensity neurosurgical scenarios. The entire system non-invasively acquires synchronous images from a microscope, requiring no modification to existing surgical equipment, not obstructing the surgical field, and not intruding into sterile areas. It is highly compatible, lightweight, and balances the convenience of surgical operation with aseptic standardization. It realizes the digital, standardized, and refined upgrade of the hemostasis process in neurosurgical microsurgery, effectively improving the overall safety, standardization, and efficiency of the surgery. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of a wearable smart glasses system provided by the present invention.
[0020] Figure 2 This is a schematic flowchart of a control method for wearable smart glasses provided by the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] This invention discloses a wearable smart glasses system, such as... Figure 1 As shown, it includes: the smart glasses body, an image acquisition module, an image preprocessing module, a reference object recognition module, a size calculation module, and a display module; The smart glasses integrate a high-definition miniature camera and a near-eye display optical module for capturing microscope images; The image acquisition module is used to acquire microscope images in real time; The image preprocessing module is used to preprocess the microscope image; The reference object recognition module has a built-in trained medical image recognition model, which is used to recognize various standard reference objects of known fixed size in the surgical field based on preprocessed microscope images. The size calculation module is used to establish a mapping relationship between the image pixel size and the actual physical size of the standard reference object, calculate the actual area and boundary size of the bleeding area and tissue defect area in the surgical area, and automatically convert the appropriate surgical consumables specifications by combining the matching algorithm. The display module is used to project the calculated specifications of surgical consumables onto the glasses lens in real time in the form of semi-transparent AR floating labels. It adopts information floating display, does not obstruct the nurse's naked eye vision, and realizes imperceptible visual prompts during the operation.
[0023] The smart glasses themselves are lightweight AR glasses that integrate a high-definition miniature camera and a light-transmitting near-eye display. They can be worn stably and are suitable for long-term surgical scenarios without obstructing the surgical area or affecting the doctor's operation, providing instrument nurses with an independent and interference-free AR visualization display field.
[0024] The image acquisition module, relying on the miniature camera mounted on the glasses, captures the output screen of the microscope display in real time and synchronously acquires images of the surgical field during the operation, rather than directly photographing the patient's surgical area. This avoids the problem of equipment intruding into the sterile area and obscuring the surgical field, thus preventing contamination and ensuring that the acquired image is completely synchronized with the surgeon's microscope field of view in real time.
[0025] The image preprocessing module uses a high-performance edge computing chip to complete image preprocessing without relying on cloud devices, eliminating network latency and meeting the requirements for millisecond-level real-time response during surgery.
[0026] The reference object recognition module has a built-in trained medical image recognition model that can automatically identify various reference objects of known fixed size in the surgical field, including three categories: standard surgical instruments, routine anatomical structures, and intraoperative aseptic markers, providing an objective and accurate quantitative benchmark for wound size conversion.
[0027] The size calculation module establishes a mapping relationship between image pixel size and real physical size based on the identified standard reference object, accurately calculates the actual area and boundary size of the bleeding area and tissue defect area in the surgical area, and automatically converts the length and width specifications and the quantity of the matching brain cotton pads and gelatin sponges to be used in combination with the matching algorithm.
[0028] The display module projects core information such as recommended specifications, quantity, cutting reference lines, and replacement countdown of consumables obtained by the system into the glasses lenses in real time as semi-transparent AR floating markers. The information is displayed floatingly without obstructing the nurse's field of vision, achieving seamless visual prompts during the operation.
[0029] It also includes a communication module, equipped with a low-latency wireless communication unit, which enables real-time data interaction between the smart glasses and the microscope display device and the system's local terminal. It supports dynamic synchronization of surgical field parameters, real-time retrieval of the consumables database, and immediate transmission of abnormal prompt signals, ensuring that the entire system operates stably and with low latency for a long time.
[0030] Specifically, the image preprocessing module includes a pre-denoising unit, a perspective distortion correction unit, an RGB-HSI conversion unit, a hybrid threshold segmentation unit, a grayscale conversion unit, a multi-resolution feature decomposition unit, and a feature vectorization unit connected in sequence. The multi-resolution feature decomposition unit is used to decompose a grayscale image into a high-resolution feature map HR and a low-resolution feature map LR. The low-resolution feature map (LR) is scaled using bilinear interpolation to obtain a size-matched medium-resolution feature map (MR), and gradient texture detail features are extracted from the medium-resolution feature map (MR). ; A 3×3 Sobel gradient operator is used for MR filtering to extract gradient texture detail features. ; ; In the formula, Coordinates of the medium-resolution feature map The pixel grayscale value; The grayscale values of the original pixels in the neighborhood of the low-resolution feature map; These are the bilinear interpolation weighting coefficients, dynamically allocated based on the pixel Euclidean distance, with a value range of [0,1]. The offset of neighboring pixels; Based on the grayscale and luminance components of the high-resolution feature map HR Gray-scale luminance components of the medium-resolution feature map (MR) High-frequency detail features of the image were extracted. ; The feature vectorization unit is used to extract high-frequency detail features of the image using a uniform fixed block size and overlap parameters. Gradient texture detail features The feature blocks are divided into blocks and converted into high-dimensional image vectors X and low-dimensional image vectors Y, respectively, as the output standardized feature vectors to eliminate feature shifts caused by illumination and magnification.
[0031] Among them, high-frequency detail features of the image ; Set a uniform block size of 64×64 and an image overlap of 16 pixels. , Synchronous block processing.
[0032] The pre-denoising unit includes: before grayscale conversion, using a 5×5 two-dimensional median filter to denoise the original RGBA image. The two-dimensional median filter calculation formula is as follows: ; in, Output image coordinates Pixel values after filtering; The original pixel grayscale values within the 5×5 filtering window; Solving for the median operator; This represents the row and column offset within the window, with a value range of [-2, 2].
[0033] The perspective distortion correction unit includes: extracting the coordinates of the four corners of the microscope screen rectangle. As vertices of a distorted image, the vertices of a standard distortion-free rectangle Construct the homography matrix T: ; The pixel perspective forward mapping equations are: ; in, These are the pixel coordinates of the distorted image; These are the corrected standard pixel coordinates; The mapping coefficients of the homography matrix are to be determined; the system of equations represents the perspective transformation mapping relationship, used to correct image distortion caused by shooting tilt.
[0034] The entire pixel of the correction image is traversed, and the corresponding original pixel value is obtained by reverse mapping through one-way inverse transformation. After correction, the invalid calibration area at the edge of the image is cropped, and only the effective image of the surgical field is retained.
[0035] The RGB-HSI conversion unit is used for RGB to HSI color space conversion, and the complete set of equations is as follows: ; in, Represents the hue component, characterizes the type of pixel color, is in rad, and has a value range of [0, 2π]. The red wound area is concentrated in the range of [0, π / 3]. It represents the color angle in the RGB three-dimensional vector space. It is an intermediate calculation parameter in rad, which is solved by the inverse cosine function and is used to distinguish between red / green / blue color systems. Saturation component represents the intensity of color, is dimensionless, and has a normalized value range of [0,1]; S approaches 1 for blood regions and approaches 0 for gray-white normal tissues. represents the brightness component, characterizing the brightness of a pixel. It is dimensionless, with a value range of [0, 255], and is directly affected by the brightness of the microscope light source. Represents the grayscale values of the red, green, and blue channels of the original image. It is an integer type and its value range is [0, 255]. This represents the inverse cosine trigonometric function operator, with the output range limited to [0,π]. This represents the operator for solving the minimum pixel value of the three channels; by stripping away the illumination interference of the luminance component I, precise pixel-level segmentation of the bleeding wound and normal tissue is achieved based on the high-discrimination saturation component S.
[0036] Specifically, the hybrid threshold segmentation unit employs a saturation-based dual-threshold weighted fusion algorithm, as shown in the following formula: ; ; in, This represents the comprehensive adaptive saturation threshold used for the final binary segmentation of the wound surface. It is dimensionless and ranges from [0,1]. This indicates that the Otsu method is an adaptive segmentation threshold, which is solved by maximizing the inter-class variance between foreground and background, and has strong resistance to overall illumination shift. The minimum error segmentation threshold is determined by minimizing the pixel classification misclassification error and has strong resistance to local reflection noise. This indicates that the weighted fusion has a fixed weight hyperparameter, with a fixed value of 0.4 in the engineering process. The algorithm prioritizes the global segmentation effect of the Dajin method while also taking into account the local anti-interference ability of the minimum error method. This represents the fusion weight corresponding to the minimum error threshold, with a value of 0.6; The binary segmentation determination rule is based on pixel saturation. Pixels identified as foreground pixels of bleeding / wound are marked as mask ground truth value 1; Pixels identified as surgical instruments or normal tissue backgrounds are marked as mask ground truth value 0, and a binary mask image is output.
[0037] The grayscale conversion unit converts the corrected and segmented RGB image into a single-channel grayscale image. .
[0038] Specifically, the medical image recognition model adopts an improved NAM-YOLACT algorithm, targeting microsurgical scenarios with low signal-to-noise ratio, blood occlusion, and illumination distortion. It employs an optimized one-stage lightweight instance segmentation network, abandons the redundant convolutional structure of the original YOLACT, and integrates NAM normalized attention mechanism, medical multi-scale feature enhancement, and edge sparse compression strategy. It is designed for millisecond-level inference of embedded edge NPU, including: input layer, lightweight residual backbone layer, NAM normalized attention enhancement layer, multi-scale FPN fusion layer, dual-branch inference layer, and processing output layer; The input layer receives the preprocessed microscope image, including the standardized feature vector output by the image preprocessing module and the RGB microscopic image. The input layer has a built-in dimension normalization interface, which receives three types of auxiliary information: overlapping block features, real-time microscope magnification labels, and HSI saturation masks. It removes invalid background feature components and unifies the input tensor dimension to [3, 1080, 1920] to avoid network input distribution drift caused by brightness shift of the microscopic image. The lightweight residual backbone layer adopts the ResNet-50-DW bottleneck backbone network, replaces all standard 3D convolutions with depthwise separable convolutions, removes the deep redundant downsampling layers and global average pooling structure of the original network, and sets up a four-level residual stage to extract image features step by step, extracting surface texture features, mid-level geometric contour features and deep semantic features of the surgical field respectively. Compared with the original ResNet-50, the number of parameters is compressed by 32%, which reduces the amount of floating-point operation while retaining the fine-grained features of millimeter-level reference objects, and is suitable for low computing power deployment of edge devices. The NAM normalized attention enhancement layer is embedded in the output of each level of the lightweight residual backbone layer and the back end of the multi-scale FPN fusion layer. It includes two parallel branches: channel attention and spatial attention. The channel attention branch relies on batch normalization to correct the feature channel distribution and combines L1 regularization constraints to strengthen the weight of effective feature channels. The spatial branch performs dynamic recalibration of the weights of the pixel region in the surgical field. Together, they suppress invalid feature responses from strong light reflections, blank backgrounds, and blood areas during surgery, accurately focus on the target area of the standard reference object, and solve the problems of overactivation and false response of the native attention mechanism in medical scenarios. The multi-scale FPN fusion layer constructs a five-level feature pyramid structure from P3 to P7, optimizing the upsampling and downsampling ratios for the size span of the microscopic reference object. The P3 level is adapted to 4mm micro-particle calibration blocks, the P4-P5 levels are adapted to 5-10mm conventional calibration instruments, and the P6-P7 levels are adapted to large-scale surgical instruments. The fusion process embeds high-frequency detail features and gradient texture features output from the image preprocessing module to complete the edge information of the micro-reference object and eliminate the defect of small target feature loss caused by deep feature downsampling. The dual-branch inference layer includes two parallel sub-branches: a detection and regression branch and a prototype mask generation branch. It features a one-stage ROI-free pooling architecture. The detection and regression branch outputs the target class probability, bounding box offset, and global matching confidence. The prototype mask generation branch generates 32 sets of medical prototype mask basis vectors and generates refined instance masks through feature weighted combination. The entire process is free of candidate box pooling operations, avoiding the loss of small calibration block features and ensuring 30fps real-time inference performance. The processing output layer integrates a triple post-processing mechanism of Harris-LACC joint mismatch elimination, confidence screening, and medical adaptive nonmaximum suppression. It eliminates mismatch feature pairs caused by blood and light interference by normalizing the correlation coefficient through LACC, filters invalid detection instances with confidence less than 0.85, removes overlapping detection boxes, and finally outputs the reference object category, subpixel-level bounding box coordinates, and the total number of effective coverage pixels M of the reference object in a structured output, which is pushed to the size calculation module to complete the pixel-physical size mapping calculation.
[0039] Perform pixel traversal statistics on the filtered valid reference object binary mask: only count the foreground pixels marked as true value 1 in the mask, and automatically remove background pixels such as mask holes, surrounding surgical instruments, and tissues to complete the pixel count of the pure reference object area.
[0040] The total number of pixels M obtained from the statistics is used as a structured output parameter, and is pushed to the size calculation module along with the reference object category and sub-pixel boundary coordinates for subsequent physical mapping solution of pixel equivalent Spixel.
[0041] This algorithm, based on the native YOLACT instance segmentation network, introduces a NAM normalized attention enhancement layer to replace the traditional CBAM and SE attention for microsurgery scenarios. It leverages batch normalization to naturally suppress intraoperative illumination distribution shifts, and combines L1 regularized attention weight sparsity constraints to force the network to weaken the activation of large areas of blood and blank tissue background features, while strengthening the feature response of small calibration blocks. The backbone network is modified into a depthwise separable residual structure, and the K-SVD sparse dictionary compression algorithm is used to sparsify the parameters of the fully connected layers and mask prototype vectors, resulting in an overall model parameter compression of 60%, making it suitable for low-computing-power NPUs at the edge of glasses. Furthermore, medical prior knowledge is embedded, with HSI saturation wound masks and microscope magnification parameters embedded into the network input layer and loss function layer, enabling the network to distinguish between blood foregrounds, standard references, and normal tissue, thereby fundamentally reducing the rate of missed and false detections in occluded scenes.
[0042] Specifically, the K-SVD sparse dictionary compression algorithm is expressed as: ; The complete formula for fitting the high-resolution dictionary Cholesky decomposition is: ; in, This represents a low-resolution feature dictionary, with dictionary atoms adapted to the low-dimensional feature representation of embedded devices. This represents a high-resolution feature dictionary used to reconstruct refined surgical feature information; Represents the feature sparse coefficient vector, the core compression parameter, and the projection weights of the features onto dictionary atoms; This represents a sparse constraint condition. The L0 norm (number of non-zero elements) is used, and K=16 is the number of the largest non-zero coefficients, which forces feature sparsity. This represents the low-dimensional surgical feature vector to be reconstructed from the network output; The squared L2 norm represents the feature reconstruction error; the smaller the error, the higher the dictionary fitting accuracy. This represents the KL divergence balance hyperparameter, with a fixed value of 0.15, which is the weight of the balance reconstruction error and the difference in feature distribution. The KL divergence (relative entropy) operator measures the distribution of features learned by the model. With the optimal target distribution Differences; This represents the original high-dimensional surgical feature matrix, which contains all refined surgical field features of the training set. This represents the lower triangular invertible matrix obtained from the Cholesky decomposition, ensuring the numerical stability of matrix inversion and avoiding singularity in the decomposition. It represents the minimum value solving operator, which traverses the dictionary and sparse coefficients to find the globally optimal combination; it is used to complete the sparse compression of model parameters, reducing the overall parameters by 60%, and adapting to low computing power real-time inference of edge NPU while retaining recognition accuracy.
[0043] Specifically, the end-to-end recognition and processing flow of the model
[0044] S301, Receive the normalized feature vector, RGB original image, and saturation binary mask output by the image preprocessing module; S302, a lightweight ResNet-50-DW backbone network extracts texture, geometric and semantic features at three levels step by step to generate four-level basic feature maps; S303. Features at all levels are fed into the NAM module to perform channel dimension weight normalization and spatial dimension region focusing, respectively, to suppress invalid feature responses. S304 and FPN pyramid fusion multi-level enhancement features, embedding high-frequency texture features to complete the edge details of small targets, and generating a five-level multi-scale feature map; S305, Detection branch predicts category and bounding box, mask branch generates medical prototype mask and weighted combination, outputs initial detection result; S306. Use Harris corner matching and LACC correlation coefficient to remove false matches of light / blood, filter invalid targets with confidence <0.85, and remove overlapping detection boxes; S307 outputs the category, boundary coordinates, instance mask, and total number of covered pixels M of the valid reference object, providing accurate benchmark parameters for subsequent pixel equivalent calculation and wound size measurement.
[0045] Specifically, the Harris corner neighborhood gray-level similarity complete matching formula is as follows: ; in, Represents feature points in real-time surgical field images Feature points of the standard reference template The neighborhood gray-level similarity is dimensionless; This represents the Harris corner pixel coordinates extracted from the real-time image; This represents the Harris corner pixel coordinates of the built-in reference template image; This represents a fixed 5×5 pixel circular matching neighborhood window, which limits the local calculation range of feature matching and suppresses global interference. Representing the neighborhood The total number of valid pixels within the window is used for similarity mean normalization to eliminate the influence of window size. This represents the grayscale value of each pixel within the neighborhood of the real-time surgical field image; This represents the grayscale value of each pixel within the neighborhood of the built-in standard reference template image; This represents the arithmetic mean of pixel grayscale values within the matching neighborhood of a real-time image. This represents the arithmetic mean of pixel grayscale values within the matching neighborhood of the template image; by performing a mean-free cross-correlation operation, the interference of global illumination offset on feature similarity calculation is eliminated.
[0046] The complete formula for determining the LACC normalized correlation coefficient is: ; in, represents the normalized correlation coefficient, which is dimensionless and has a standard value range of [-1, 1]. The closer the value is to 1, the higher the confidence of the feature point matching; a value close to -1 indicates a negative correlation matching. This represents the neighborhood gray-level cross-correlation similarity obtained from the aforementioned Harris matching; It represents the pixel gray-level variance of the matching neighborhood of a real-time image, and characterizes the degree of local gray-level dispersion of the real-time image; It represents the pixel gray-level variance of the neighborhood of the reference template matching, and characterizes the degree of local gray-level dispersion of the template image; Matching and filtering rules: The system presets a fixed confidence threshold of 0.7, retaining only... Feature point pairs are considered valid matches, while the rest are deemed false matches due to blood occlusion or lighting interference and are directly discarded.
[0047] Supervised training was performed using a total loss function with L1 regularization. The dataset consisted of 1500 microscopic annotations including strong light reflections, blood partial occlusion, and different magnifications, supplemented by Mosaic medical image augmentation samples. The NAM attention weights were iteratively updated via backpropagation. By combining the parameters of the backbone network, the medical inference weights are finally obtained through convergence.
[0048] Specifically, in the NAM normalized attention enhancement layer, the channel attention output is: ; Spatial attention output: ; Channel / spatial weight normalization allocation formula: ; in, This represents the channel attention weight matrix, with dimensions consistent with the number of channels C in the feature map. It is used to strengthen the weights of effective feature channels of the reference object and suppress ineffective channels. This represents the spatial attention weight matrix, with dimensions consistent with the feature map spatial dimensions (H,W), used to focus the target region of the surgical field and suppress the response of the background region; This represents an S-shaped nonlinear activation function that normalizes the input weights to the interval [0,1], thus adapting to the probability distribution characteristics of the attention weights. This represents the batch normalization operator, which eliminates the feature distribution offset between batches, accelerates model convergence, and improves robustness. The network intermediate layer feature maps representing the input to the attention module correspond to the channel branch and spatial branch inputs, respectively. This represents the global pixel mean of all feature maps within a single training batch; This represents the global variance of pixels across all feature maps within a single training batch. This represents the normalized learnable scaling factor, which is iteratively updated by backpropagation of the model to adjust the magnitude of the feature distribution. This represents the normalized learnable offset coefficient, which is iteratively updated by backpropagation of the model to adjust the center position of the feature distribution. This represents a numerically stable minimum constant with a fixed value in engineering applications. This prevents the calculation from having a zero denominator, thus avoiding computational singularities. This represents the normalized channel and spatial attention weight coefficients; Represents the original learned weights for the i-th channel / spatial location; This represents the total number of channels and the total number of spatial locations in the feature map; It represents the Hadamard product of matrices (element-wise multiplication operator) to achieve precise weighting of weights and feature maps.
[0049] The complete function of the total loss of the model with L1 regularization is: ; The formula for decomposing the basic loss during detection is as follows: ; in, This represents the global total loss of the NAM-YOLACT model. It is dimensionless and used for backpropagation to update all network parameters. The basic loss for medical instance segmentation consists of three components: classification loss, bounding box regression loss, and instance mask segmentation loss. This represents the target classification loss function, which uses cross-entropy loss to determine the category to which a feature belongs; This represents the bounding box regression loss function, which uses CIoU loss to optimize the accuracy of the detection box coordinates. The instance mask segmentation loss function is represented by a binary cross-entropy loss, which accurately segments the outline of the reference object. The standardized feature vector representing the network input is output by the preprocessing module; The model represents the input features The output includes classification prediction, bounding box coordinate prediction, and mask probability prediction values. This represents the manually labeled actual classification labels, actual bounding box coordinates, and actual mask binary labels. This represents the L1 canonical equilibrium hyperparameter, with fixed values in the engineering process. Adjust the sparsity of the weights; This represents the L1 norm operator, which imposes a sparsity constraint on the NAM attention coefficients, forcibly suppressing the weights of ineffective background features in the surgical field. This indicates that the NAM module can learn channels and the original weight coefficients of the space, which correspond one-to-one with the parameters of the aforementioned attention formula.
[0050] Specifically, the size calculation module establishes a pixel-to-real physical size mapping based on the identified standard reference object, uses eight-neighborhood tracking of the wound contour to solve for the bleeding / defect area, and graded matching of surgical consumables.
[0051] Specifically, based on the microscope's real-time magnification factor k, the inherent reference length and width of the standard reference object are compared with those at the factory. , Magnification correction is performed to obtain the equivalent true length and width of the reference object adapted to the current microscope field of view. , , Substituting the total number of pixels M effectively covered by the reference object into the pixel equivalent formula Calculate the actual physical area corresponding to a single pixel. ; A dynamic pixel-to-millimeter mapping relationship is established, and parameters are updated in real time when the microscope magnification is switched during surgery to eliminate measurement system errors caused by zooming. After the mapping benchmark is completed, the complete edge of the wound is extracted using an eight-neighbor counterclockwise contour tracking algorithm based on the binary mask of the bleeding wound output by the image preprocessing module. The invalid pseudo-boundaries formed by breakpoints, noise, and occlusion are filtered out by an eight-direction vector encoding judgment rule, retaining only the valid contour points. The pure effective pixels inside the wound are counted. All pixels of the outline boundary ; To address the inherent statistical bias that half of the boundary pixels cover the wound and the other half belong to the external background, a half-area correction formula is used. Calculate the equivalent total pixels, and then combine them with the already solved actual physical area. The measured area of the wound was obtained: ; At the same time, the actual physical area The square root yields the actual side length coefficient of a single pixel. The actual maximum length and width of the wound are calculated by combining the extreme pixel coordinates of the wound outline. This boundary correction mechanism can control the relative error of area measurement within ±7.5%, meeting the accuracy requirements for clinical use in microsurgery. Based on the actual measured area and dimensions of the wound, it automatically matches the corresponding specifications of gelatin sponge, brain cotton pads, and estimated usage by retrieving locally fixed neurosurgical clinical grading thresholds (3mm², 10mm²). Then, a fixed safety redundancy coefficient ξ=1.1 is introduced, and a proportional conversion formula is used. Calculate the cutting dimensions of consumables with a 10% coverage margin, and finally send the wound length and width, actual area, consumable model, estimated usage, and consumable cutting reference dimensions to the display module.
[0052] The system drives four layers of AR to render numerical text, consumable labels, and auxiliary clipping dotted lines, providing instrument nurses with intuitive reference during surgery in a semi-transparent floating form. This enables integrated real-time output of wound quantification and intelligent recommendation of surgical consumables.
[0053] Specifically, the formula for calculating the physical pixel equivalent is: ; The formula for dynamic correction of microscope magnification is: ; in, Represents pixel equivalent, a core physical mapping parameter, characterizing the actual physical area of the surgical region corresponding to a single image pixel; Indicates the actual length and width of the reference standard as specified by the manufacturer, representing fixed physical quantities; This indicates the corrected equivalent length and width of the reference object at the current real-time magnification of the microscope; This represents the real-time magnification factor of the microscope, with a value range of [0.5, 20]. This represents the total number of pixels covered by the effective standard reference in the image, counting only the effective pixels within the mask range and removing background interference pixels. Real-time updates when switching microscope magnification Dynamically refresh pixel equivalent Completely eliminate the system error in size measurement caused by zoom; Specifically, the formula for accumulating pixels within the eight-neighbor contour is as follows: ; Complete logical formula for determining the validity of eight-directional contour points: ; in, This represents the total number of purely valid pixels inside the closed contour of the wound, an integer type, excluding boundary pixels and hole pixels; This represents the total number of boundary points of the current wound closure contour, and the upper limit of the summation operation. The horizontal column coordinates of the image representing the i-th and i+1-th adjacent boundary points on the contour, in pixels; This indicates that the larger operator is used to avoid invalid pixel statistics caused by negative coordinate differences; This indicates the Boolean determination result for the validity of boundary points. Valid points are included in the area calculation, while invalid points are determined to be noise or pseudo-boundaries. This represents the forward vector encoding of the current boundary point in eight directions, with integer values [0,7] representing the forward direction of the contour within the eight neighborhoods; This represents the vector code of the current boundary point in eight directions, with integer values [0,7] representing the contour backtracking direction; A dedicated invalid boundary identifier represents contour breakpoints, isolated noise, and pseudo-boundaries at occluded breaks; it accurately filters valid contour points, avoiding pixel statistical errors caused by wound holes, blood noise, and contour breaks.
[0054] Specifically, the total pixel half-area correction of the wound (half of the boundary pixels are outside the tissue) includes: The formula for correcting the half-area of boundary pixels is: ; The formula for calculating the final actual area of the wound is: ; The formula for converting the length and width of a wound is: ; The intermediate formula for converting pixel side length is: ; in, This represents the total number of all valid pixels on the closed contour boundary; This represents the equivalent total number of pixels in the wound area after half-area correction; the principle is that half of the boundary pixels span the inside of the wound and half span the outside background, so half of the area is used for weighting. This indicates the final measured effective area of bleeding / soft tissue defect in the surgical area; This represents the actual side length coefficient corresponding to a single pixel, which is derived from the square root of the pixel equivalent. Represents the horizontal column coordinates of the wound outline, specifically the maximum and minimum pixel column coordinates. Represents the vertical row coordinates of the wound outline, specifically the maximum and minimum pixel row coordinates. This indicates the actual maximum longitudinal length of the wound, in mm; This indicates the actual maximum horizontal width of the wound. This corrective scheme solves the problem of boundary bias in traditional pixel statistics, controlling the relative error of area calculation within ±7.5%, which meets the requirements of clinical accuracy.
[0055] The formula for determining the matching of consumable grades based on wound area is as follows: ; The formula for proportional conversion of material cutting dimensions is as follows: ; in, It represents the actual measured effective area of the wound and is the core criterion for graded matching. Indicates the measured maximum length and maximum width of the wound; The recommended length and width dimensions for materials labeled in the AR layer are indicated. This represents the hemostasis safety redundancy coefficient. The engineering solidification value is 1.1, reserving a 10% edge coverage margin to prevent bleeding from the wound edge and incomplete coverage. Interval threshold explanation: 3mm² and 10mm² are clinically calibrated thresholds for neurosurgical microsurgery, which are based on four reference patent medical databases; The matching results, consumable usage, and cutting size parameters are pushed to the display module in real time, generating a visual AR-assisted cutting dotted line for instrument nurses to refer to directly during the operation.
[0056] In one specific embodiment of the present invention, such as Figure 2 As shown, a control method for wearable smart glasses includes: real-time acquisition of microscope images; The microscope images are preprocessed; Identify various standard reference objects of known fixed size within the surgical field based on preprocessed microscope images; A mapping relationship between image pixel size and real physical size is established for the standard reference object. The actual area and boundary size of the bleeding area and tissue defect area in the surgical area are calculated. The appropriate surgical consumable specifications are automatically converted by the matching algorithm. The calculated specifications of surgical consumables are projected onto the glasses lenses for display in real time.
[0057] In one specific embodiment of the present invention, a complete clinical embodiment of a wearable smart glasses system is provided, taking a neurosurgical brain tumor microsurgical resection scenario as an example: Surgical type: Microsurgical resection of frontal lobe glioma. During the operation, there was bleeding in the brain tissue and local tissue defects, requiring continuous use of brain cotton pads and gelatin sponges for hemostasis. Equipment configuration: Existing surgical microscope (supporting 0.5~20x stepless zoom), smart glasses of the present invention, worn by scrub nurses, with built-in edge computing NPU, miniature camera, light-transmitting near-eye display module, 5mm standard micro hemostatic forceps as the surgical standard reference, and 3mm square calibration block for intraoperative use; The surgeon operates the microscope, and the scrub nurse wears smart glasses and is responsible for delivering hemostatic supplies.
[0058] This invention, through synchronous acquisition of microscope images, multi-step image preprocessing, identification of surgical field standard reference objects, establishment of pixel-to-actual size conversion relationship, automatic calculation of bleeding / defect wound size, matching and recommendation of hemostatic consumables, and real-time visualization of results, achieves full edge local computation without cloud requirements and millisecond-level real-time response.
[0059] The specific steps are as follows: Step 1: The scrub nurse puts on the lightweight AR smart glasses, starts the built-in system of the smart glasses, completes the device self-test, and initializes the camera, near-eye display optical module, edge computing chip, and wireless communication module. Adjust the angle of the miniature camera on the glasses to align with the external display screen of the operating room microscope, ensuring that the camera fully captures the surgical field output by the microscope, with no obstruction and no severe perspective distortion. The image acquisition module is activated to establish a real-time image acquisition link, continuously and synchronously capturing the surgical field images output by the microscope throughout the entire surgery, which are completely synchronized with the surgeon's microscope field of view.
[0060] Step 2: After acquiring the original microscopic images, the image preprocessing module completes 7 layers of image processing in a fixed order to eliminate interference from lighting, reflection, distortion, and blood, providing standardized input for intelligent recognition. Pre-processing noise reduction: The original image is processed using a 5×5 median filter to filter out image noise caused by microscope light reflection and blood particles; Perspective distortion correction: Identify the four vertices of the microscope screen, construct a homography transformation matrix, correct the image stretching and distortion caused by the camera's tilt shooting, and crop out invalid images outside the screen; RGB-HSI color space conversion: Converting RGB color images to the HSI color model, removing brightness and illumination interference, and extracting the saturation channel separately, the saturation value of blood wounds is significantly higher than that of normal brain tissue, greatly improving the distinction. Hybrid saturation threshold segmentation: The comprehensive segmentation threshold is calculated by combining the Otsu method and the minimum error method with dual threshold weighting. It automatically distinguishes between bleeding wounds (foreground) and normal tissues and surgical instruments (background), and outputs a binary mask image of the wound, with white areas representing bleeding defect areas. Grayscale conversion: Converts the corrected and segmented color image into a single-channel grayscale image to reduce the amount of subsequent computation; Multi-resolution feature decomposition: The grayscale image is split into two feature maps: high resolution and low resolution. The low resolution image is interpolated and scaled to obtain a medium resolution feature map. Gradient texture details and high-frequency details of the image are extracted respectively. Feature vectorization and standardization: The two types of detailed features are processed in blocks and converted into high-dimensional and low-dimensional feature vectors in a unified format. This eliminates feature offsets caused by different microscope magnifications and illumination, and outputs standardized features to the reference object recognition module.
[0061] Step 3: Simultaneously feed the preprocessed standardized image and feature vectors into the built-in trained medical image recognition model for real-time inference and recognition at the local edge. The input layer receives the original microscopic image, standardized feature vector, wound saturation mask, and real-time microscope magnification parameters, and unifies the image tensor dimension. A lightweight ResNet-50-DW backbone network is used to extract three levels of image features from the surgical field: surface texture, mid-level contour, and deep semantics. The NAM normalized attention enhancement layer is embedded in the feature output of each level, and opens the channel and spatial attention branches in parallel. It automatically weakens the invalid features of large areas of blood and blank brain tissue, and focuses on standard reference objects with known size in the picture. Multi-scale FPN five-level feature pyramid fusion of multi-level enhanced features completes the edge details of small calibration blocks and avoids feature loss during downsampling of small reference objects; The detection branch predicts the target category and bounding box through dual-branch parallel inference, while the masking branch generates a reference instance segmentation mask. Post-processing to filter interference: False matching targets caused by blood and reflection are removed by Harris corner matching and LAC correlation coefficient, invalid detection boxes with confidence scores below 0.85 are filtered out, and overlapping recognition results are removed; Structured output of valid reference object information: output reference object type (5mm hemostat / 3mm calibration block), subpixel-level precise contour coordinates, and the total number of valid pixels M covered by the reference object in the image, which are then passed to the size calculation module.
[0062] Step 4: Establish a dynamic size conversion benchmark based on the identified standard reference objects, automatically calculate wound parameters and match hemostatic consumable specifications; Specifically, the microscope reads the real-time magnification factor, corrects the fixed length and width of the standard reference object at the factory, and calculates the actual physical area corresponding to a single pixel based on the total number of pixels of the reference object. When the microscope zooms, this parameter is automatically updated in real time to eliminate the size calculation error caused by zooming. The eight-neighbor counterclockwise contour tracking algorithm is invoked to extract the complete edges of bleeding and tissue defects based on the preprocessed output of the wound binary mask; breakpoints and noise pseudo-boundaries are filtered out by eight-direction vector encoding, and only valid contour points are retained; the pure valid pixels inside the wound and the pixels of the contour boundary are counted respectively. A boundary half-area correction algorithm is used to correct the statistical deviation of boundary pixels, calculate the equivalent total pixels of the wound, and combine the pixel equivalent to obtain the true physical area of the wound; then the actual side length of a single pixel is converted, and combined with the extreme coordinates of the contour pixels, the actual maximum length and width of the wound are calculated; thus meeting the clinical quantitative needs. Retrieve local clinical grading thresholds for solidified neurosurgery, automatically match the corresponding type of gelatin sponge and brain cotton pads based on the wound area, and output the estimated quantity to be used. By introducing a safety redundancy factor, the recommended material cutting size is obtained by proportionally enlarging the actual length and width of the wound and reserving a 10% coverage margin. The actual length and width of the wound, the measured area, the recommended consumable model, the estimated usage, and the reference dimensions for consumable cutting are all sent to the display module.
[0063] Step 5: The display module receives all parameters output by the size calculation module and renders four semi-transparent AR layers to display the wound size numerical text, consumable model identification, estimated consumable usage, and cutting auxiliary dotted lines; All information is projected onto the light-transmitting lens of the smart glasses in a floating, semi-transparent form. The transparency of the annotation layer is adjustable and will not obstruct the visual field of the scrub nurse to observe the operating table and instruments. The scrub nurse wearing glasses can simultaneously and intuitively see the quantitative data of the surgical field and the recommendations of consumables, without the need for the surgeon to verbally announce the wound size and consumable model, and without having to bend down and repeatedly compare the actual size; If the microscope magnification is switched or the surgical field wound is enlarged / reduced, the entire system repeats steps 2-4 in real time, refreshing the displayed content and updating dynamically throughout the process.
[0064] Traditionally, the surgeon verbally describes the wound size and vaguely indicates small / large size head cotton, which can easily lead to nurses misunderstanding and picking the wrong item. This embodiment uses AR to directly provide quantitative dimensions and consumable models, eliminating information transmission errors. Nurses no longer need to frequently bend down to check the instrument table and compare consumables, reducing cross-contact between sterile and non-sterile areas and lowering the risk of postoperative infection. The system automatically calculates and pushes consumable plans in milliseconds, saving the time spent communicating with medical staff and repeatedly confirming consumables, and shortening the time spent on intraoperative hemostasis. By relying on standard reference objects to establish objective size conversion benchmarks, the subjective visual estimation of doctors is eliminated, and the size of the wound and the selection of consumables have unified quantitative standards, making the surgical procedure more standardized. The microscope display screen is photographed only through the glasses' camera, without requiring modification of the existing surgical microscope or intrusion into the sterile surgical field. It is lightweight and can be worn for extended periods of surgery without significant weight-bearing.
[0065] After tumor resection and hemostasis are completed, the image acquisition module is turned off to stop real-time acquisition of the surgical field; the cached data of the surgical images and wound measurement in the edge computing chip is cleared; the display layer is turned off and the system inference program is exited; the smart glasses are removed, and the surface is disinfected and stored, ready for reuse in the next surgery.
[0066] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0067] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A wearable smart glasses system, characterized in that, include: The smart glasses consist of a main body, an image acquisition module, an image preprocessing module, a reference object recognition module, a size calculation module, and a display module. The smart glasses integrate a camera and a near-eye display optical module for capturing microscope images; The image acquisition module is used to acquire microscope images in real time; The image preprocessing module is used to preprocess the microscope image; The reference object recognition module has a built-in trained medical image recognition model, which is used to recognize various standard reference objects of known fixed size in the surgical field based on preprocessed microscope images. The size calculation module is used to establish a mapping relationship between the image pixel size and the actual physical size of the standard reference object, calculate the actual area and boundary size of the bleeding area and tissue defect area in the surgical area, and automatically convert the appropriate surgical consumables specifications by combining the matching algorithm. The display module is used to project the calculated specifications of surgical consumables onto the eyeglass lenses for display in real time.
2. The wearable smart glasses system according to claim 1, characterized in that, The image preprocessing module includes a pre-denoising unit, a perspective distortion correction unit, an RGB-HSI conversion unit, a hybrid threshold segmentation unit, a grayscale conversion unit, a multi-resolution feature decomposition unit, and a feature vectorization unit connected in sequence. The multi-resolution feature decomposition unit is used to decompose a grayscale image into a high-resolution feature map HR and a low-resolution feature map LR. The low-resolution feature map (LR) is scaled using bilinear interpolation to obtain a size-matched medium-resolution feature map (MR), and gradient texture detail features are extracted from the medium-resolution feature map (MR). ; Based on the grayscale and luminance components of the high-resolution feature map HR Gray-scale luminance components of the medium-resolution feature map (MR) High-frequency detail features of the image were extracted. ; The feature vectorization unit is used to analyze high-frequency detail features of the image. Gradient texture detail features The feature blocks are divided into blocks and converted into high-dimensional image vectors X and low-dimensional image vectors Y, which are then used as the output standardized feature vectors.
3. The wearable smart glasses system according to claim 2, characterized in that, The hybrid threshold segmentation unit employs a saturation-based dual-threshold weighted fusion algorithm, as shown in the following formula: ; in, This represents the final adaptive saturation threshold used for binary segmentation of the wound. This represents the adaptive segmentation threshold, which is solved by maximizing the inter-class variance between foreground and background. The minimum error segmentation threshold is defined by minimizing the pixel classification misclassification error. This represents the fixed-weight hyperparameter of the weighted fusion; The binary segmentation determination rule is based on pixel saturation. Pixels identified as foreground pixels of bleeding / wound are marked as mask ground truth value 1; Pixels identified as surgical instruments or normal tissue backgrounds are marked as mask ground truth value 0, and a binary mask image is output.
4. The wearable smart glasses system according to claim 3, characterized in that, The medical image recognition model adopts an improved NAM-YOLACT algorithm, including: an input layer, a lightweight residual backbone layer, a NAM normalized attention enhancement layer, a multi-scale FPN fusion layer, a dual-branch inference layer, and a processing output layer; The input layer receives the preprocessed microscope image; The lightweight residual backbone layer adopts the ResNet-50-DW bottleneck backbone network and sets four residual stages to extract image features step by step, respectively extracting surface texture features, mid-layer geometric contour features, and deep semantic features of the surgical field. The NAM normalized attention enhancement layer is embedded in the output of each level of the lightweight residual backbone layer and the back end of the multi-scale FPN fusion layer, and includes two parallel branches: channel attention and spatial attention. The multi-scale FPN fusion layer constructs a five-level feature pyramid structure, and the fusion process embeds high-frequency detail features and gradient texture features output by the image preprocessing module. The dual-branch inference layer includes two parallel sub-branches: a detection and regression branch and a prototype mask generation branch. The detection and regression branch is used to output the target class probability, bounding box offset, and global matching confidence. The prototype mask generation branch is used to generate medical prototype mask basis vectors and generate instance masks through feature weighted combination. The processing output layer eliminates mismatched feature pairs by normalizing the correlation coefficient using LACC, and finally outputs the reference object category, sub-pixel level bounding box coordinates, and the total number of effective coverage pixels M of the reference object in a structured output.
5. A wearable smart glasses system according to claim 4, characterized in that, In the NAM-normalized attention enhancement layer, the channel attention output is: ; Spatial attention output: ; Channel / spatial weight normalization allocation formula: ; in, This represents the channel attention weight matrix; Represents the spatial attention weight matrix; Represents a sigmoid nonlinear activation function; This represents the batch normalization operator; The network intermediate layer feature maps representing the input to the attention module correspond to the channel branch and spatial branch inputs, respectively. This represents the global pixel mean of all feature maps within a single training batch; This represents the global variance of pixels across all feature maps within a single training batch. This represents the normalized learnable scaling factor; Indicates the normalized learnable bias coefficient; Represents a numerically stable minimum constant; This represents the normalized channel and spatial attention weight coefficients; Represents the original learned weights for the i-th channel / spatial location; This represents the total number of channels and the total number of spatial locations in the feature map; This represents the Hadamard product of a matrix.
6. A wearable smart glasses system according to claim 1, characterized in that, The size calculation module establishes a pixel-to-real physical size mapping based on the identified standard reference object, uses eight-neighborhood tracking of the wound contour to solve the bleeding / defect area, and graded matching of surgical consumables.
7. A wearable smart glasses system according to claim 6, characterized in that, Based on the microscope's real-time magnification factor k, the standard reference object's inherent length and width are compared with the factory-made reference standard. , Magnification correction is performed to obtain the equivalent true length and width of the reference object adapted to the current microscope field of view. , ; Substituting the total number of pixels M effectively covered by the reference object into the pixel equivalent formula Calculate the actual physical area corresponding to a single pixel. ; A dynamic pixel-to-millimeter mapping relationship is established, and parameters are updated in real time when the microscope magnification is switched during surgery. After the mapping benchmark is completed, the complete edge of the wound is extracted using an eight-neighbor counterclockwise contour tracking algorithm based on the binary mask of the bleeding wound output by the image preprocessing module. Valid contour points are retained by an eight-direction vector encoding judgment rule, and the pure effective pixels inside the wound are counted. All pixels of the outline boundary ; Using the half-area correction formula Calculate the equivalent total pixels, and then combine them with the already solved actual physical area. The measured area of the wound was obtained: ; At the same time, the actual physical area The square root yields the actual side length coefficient of a single pixel. The actual maximum length and width of the wound are calculated by combining the extreme pixel coordinates of the wound outline. Based on the measured area and dimensions of the wound, the locally solidified neurosurgical clinical grading thresholds are retrieved to automatically classify and match the corresponding specifications of gelatin sponge, brain cotton pads, and quantities. A fixed safety redundancy coefficient ξ is then introduced, and a proportional conversion formula is used. The system calculates the cutting dimensions of the consumables and finally sends the wound length and width, actual area, consumable model, estimated usage, and consumable cutting reference dimensions to the display module.
8. A control method for wearable smart glasses, using the wearable smart glasses system according to any one of claims 1-7, characterized in that, include: Real-time acquisition of microscope images; The microscope images are preprocessed; Identify various standard reference objects of known fixed size within the surgical field based on preprocessed microscope images; A mapping relationship between image pixel size and real physical size is established for the standard reference object. The actual area and boundary size of the bleeding area and tissue defect area in the surgical area are calculated. The appropriate surgical consumable specifications are automatically converted by the matching algorithm. The calculated specifications of surgical consumables are projected onto the glasses lenses for display in real time.