Neurosurgery navigation system and method based on augmented reality technology

Through augmented reality technology and multi-instrument collaborative analysis, the surgical stage can be identified in real time and the navigation view can be dynamically adjusted, which solves the problem of insufficient static and dynamic information perception in existing neurosurgery navigation systems, realizes an efficient and intuitive navigation experience, and improves the safety and accuracy of surgery.

CN120605100AInactive Publication Date: 2025-09-09BEIJING EASY TIMES DIGITAL TECH

Patent Information

Application Number
CN202510684806.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing neurosurgery navigation systems have problems such as separation of information display and surgical field of view, static information that does not change dynamically with the surgical progress, and a lack of dynamic understanding and perception of complex surgical processes, which leads to untimely or overloaded navigation information, affecting the accuracy and safety of the surgery.

Method used

Using augmented reality technology, the optical tracking system monitors the position and type of surgical instruments in real time, and uses a temporal convolutional neural network and self-attention mechanism to perform multi-instrument collaborative analysis, automatically identify the current surgical stage, and extract the target visualization parameter set from the preset context-visualization rule database to dynamically adjust the content and presentation of the AR navigation view.

Benefits of technology

The AR navigation view is highly synchronized with the surgical process, providing an information-related and intuitive surgical navigation experience, and improving the safety and accuracy of neurosurgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120605100A_ABST
    Figure CN120605100A_ABST
Patent Text Reader

Abstract

The invention relates to the field of augmented reality, and particularly discloses a neurosurgery operation navigation system and method based on an augmented reality technology, which can automatically identify the current operation stage by utilizing real-time dynamic perception and understanding of an operation instrument, and take the stage information as a driving force so as to realize navigation of the neurosurgery operation. And searching in a context-visualization rule database which is pre-constructed and stores visualization rules corresponding to different operation stages, and extracting a whole set of target visualization parameter set which is most matched with the specific stage, so that the content and presentation mode of the AR navigation view are dynamically adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of augmented reality, and more specifically, to a neurosurgery navigation system and method based on augmented reality technology. Background Art

[0002] Neurosurgery involves one of the most complex and delicate structures in the human body: the brain and nervous system. These surgeries require extremely high precision; even the slightest deviation can lead to serious complications or irreversible functional impairment. Therefore, while ensuring surgical safety and effectiveness, providing surgeons with accurate, real-time spatial positioning and navigation information has been a key focus of ongoing exploration and optimization in the field of neurosurgery. The deep integration of medicine and engineering, particularly the concept of "medicine and engineering integration," has greatly advanced neurosurgery, with surgical navigation systems being a key example. These systems are designed to align high-precision preoperative medical images with the patient's actual anatomical position during surgery. This provides surgeons with information on the relative position of lesions, vital blood vessels, and neural structures during surgery, guiding surgical instruments to the target area. This is crucial for achieving minimally invasive and precise surgery.

[0003] Currently, surgical navigation systems based on medical image registration are widely used clinically. These systems typically rely on optical or electromagnetic tracking technologies to track the position of the patient's head and surgical instruments. By registering preoperative CT, MRI, and other imaging data with reference markers on the patient's head during surgery, the systems can overlay information such as the planned path and vital structures on an external display, or provide the real-time position of the instrument tip in the image. However, these traditional navigation systems have several inherent limitations. First, information is typically displayed on a separate monitor outside the surgical field of view, requiring the surgeon to frequently shift their gaze, increasing cognitive burden and disrupting the continuity of the surgical process. Second, intraoperative "brain shift" is a common problem in neurosurgery, leading to discrepancies between preoperative images and the actual anatomy during surgery, compromising navigation accuracy. Furthermore, existing navigation systems often focus on providing static anatomical positioning information and lack a dynamic understanding and perception of the complex surgical process itself. Consequently, they are unable to intelligently adjust the presentation and content of navigation information based on the actual progress and stage of the surgery, resulting in information overload or a lack of timeliness and relevance.

[0004] Therefore, an optimized neurosurgical navigation system is desired. Summary of the Invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide a neurosurgery navigation system and method based on augmented reality technology, which can automatically identify the current surgical stage by using the real-time dynamic perception and understanding of surgical instruments, and use this stage information as a driving force to search in a context-visualization rule database that pre-constructs and stores visualization rules corresponding to different surgical stages, extract a complete set of target visualization parameter sets that best matches the specific stage, and dynamically adjust the content and presentation of the AR navigation view.

[0006] According to one aspect of the present application, a neurosurgery navigation system based on augmented reality technology is provided, comprising: Tracking data acquisition module, used to obtain the marker tracking data collected by the optical tracking system; An instrument data generation module, configured to generate a real-time position and type list of the tracked instrument within the current field of view based on the landmark tracking data; a surgical stage recognition module, configured to input the real-time position and type list of the tracked instrument in the current field of view into a surgical stage recognition model to obtain a currently estimated surgical stage; a target parameter data extraction module, configured to extract a target visualization parameter set from a context-visualization rule database based on the currently estimated surgical stage; A visualization parameter prediction module, configured to calculate a visualization parameter set used for rendering the next frame based on a comparison between the target visualization parameter set and the visualization parameter set at the current moment and in combination with a smooth transition parameter; An AR navigation view generation module is configured to perform rendering based on the visualization parameter set used for rendering the next frame to obtain an AR navigation view.

[0007] According to another aspect of the present application, a neurosurgery navigation method based on augmented reality technology is provided, comprising: Acquiring marker tracking data collected by an optical tracking system; Based on the landmark tracking data, a real-time position and type list of the tracked device in the current field of view is generated; Inputting the real-time position and type list of the tracked instrument in the current field of view into a surgical stage recognition model to obtain a current estimated surgical stage; extracting a target visualization parameter set from a context-visualization rule database based on the currently estimated surgical stage; Comparing the target visualization parameter set with the visualization parameter set at the current moment and combining the smooth transition parameter to calculate the visualization parameter set used for rendering the next frame; Rendering is performed based on the visualization parameter set used for the next frame rendering to obtain an AR navigation view.

[0008] Compared with the existing technology, the present application provides a neurosurgery navigation system and method based on augmented reality technology, which can utilize the real-time dynamic perception and understanding of surgical instruments to automatically identify the current surgical stage, and use this stage information as the driving force to search in a context-visualization rule database that has been pre-built and stored with visualization rules corresponding to different surgical stages, and extract a complete set of target visualization parameter sets that best matches the specific stage, thereby dynamically adjusting the content and presentation of the AR navigation view. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 4 is a system block diagram of a neurosurgery navigation system based on augmented reality technology according to an embodiment of the present application.

[0011] Figure 2 4 is a block diagram of a surgical stage identification module in a neurosurgery navigation system based on augmented reality technology according to an embodiment of the present application.

[0012] Figure 3 This is a block diagram of a posture timing coordination unit in a neurosurgery navigation system based on augmented reality technology according to an embodiment of the present application.

[0013] Figure 4 Flowchart of a neurosurgery navigation method based on augmented reality technology according to an embodiment of the present application.

[0014] Figure 5 Schematic diagram of data flow of a neurosurgery navigation method based on augmented reality technology according to an embodiment of the present application. DETAILED DESCRIPTION

[0015] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0016] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0017] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.

[0018] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0019] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0020] To overcome these challenges, improving the real-time, intuitive, and intelligent nature of surgical navigation systems, particularly new technologies combining medicine with virtual reality or augmented reality, has become a research hotspot. Augmented reality (AR) technology can directly overlay navigation information onto the surgeon's real-world field of view, eliminating the need to shift the surgeon's gaze, significantly improving the efficiency and intuitiveness of information acquisition. However, simplistic information overlay can lead to cluttered viewing and distract the surgeon's attention. An ideal AR navigation system should be able to provide the most relevant and concise information based on the real-time status of the surgery. The surgical process is dynamic, consisting of a series of sequential or parallel operational phases, with surgeons' information needs varying during each phase. For example, during the exposure phase, the surgeon may be more focused on the location of bones and vital vessels; during the resection phase, precise visualization of the lesion boundary and surrounding critical functional areas is required. Therefore, accurately identifying the current surgical phase in real time and dynamically adjusting visualization parameters (such as transparency, color, and displayed content) in the AR navigation view based on this information are key to achieving intelligent, adaptive navigation. Traditional navigation systems or surgical record-keeping methods often rely on preoperative planning or manual annotation by the surgeon, making it difficult to achieve real-time, automatic phase recognition and dynamic adjustment of navigation information during surgery.

[0021] Given that surgical instruments are the primary tools doctors use to perform operations, their type, location, and time-series motion patterns directly reflect the tasks the doctor is performing and the current stage of the operation. Therefore, to address the issues of traditional surgical navigation systems, where information display is separated from the surgical field of view, static information does not change dynamically with the surgical progress, and there is a lack of in-depth understanding of complex surgical behaviors, this solution proposes a new paradigm. Its core lies in no longer relying solely on static navigation information from preoperative planning, but rather through real-time monitoring and analysis of the surgical instruments used by the doctor. Because the type of surgical instruments and their real-time position and motion changes in three-dimensional space contain rich information about the doctor's current operation and the stage of the operation, this technology obtains real-time data on these instruments through a high-precision optical tracking system.

[0022] Furthermore, considering that the surgical process often involves the coordinated use of multiple instruments, simple single-instrument analysis is difficult to fully capture the complex operational intentions and stage characteristics. One of the key innovations of the technical concept of this solution is the design of a method that can perform posture time series collaborative aggregation analysis on the real-time posture change time series of multiple instruments, extracting global dynamic collaborative semantic features that reflect the overall surgical process from the scattered instrument dynamic features. This collaborative semantic feature is a highly summarized and abstract representation of the current surgical status. Subsequently, this semantically rich collaborative information is input into a pre-trained surgical stage recognition model to automatically determine the current estimated surgical stage in real time.

[0023] Once the current surgical stage is identified, the system extracts the target visualization parameter set that best matches the current stage from a pre-set context-visualization rule database, based on the surgeon's varying information needs during each stage. These parameter sets meticulously define which anatomical structures, lesions, and vital vessels should be displayed in the AR navigation view, as well as their display methods, such as transparency, color, and wireframe or solid rendering. By comparing the target visualization parameter set with the currently active visualization parameter set and calculating smooth transition parameters, the system determines the optimal visualization parameter set for the next frame of AR rendering. Finally, based on this calculated rendering parameter set for the next frame, the system generates and displays the AR navigation view. This allows navigation information to be directly overlaid on the surgeon's real-world field of view in the most intuitive and relevant manner, dynamically adjusting its content and presentation as the surgical stage changes in real time. This approach enables intelligent analysis and understanding of the surgical stage based on instrument tracking data, and then intelligently controls AR visualization using this stage information, providing a highly synchronized, relevant, and intuitive surgical navigation experience. This effectively addresses key technical issues such as information disjunction, static display, and a lack of understanding of the surgical process, thereby improving the safety of neurosurgery.

[0024] In the technical solution of the present application, a neurosurgery navigation system based on augmented reality technology is proposed. Figure 1 FIG. 1 is a system block diagram of a neurosurgery navigation system based on augmented reality technology according to an embodiment of the present application. Figure 1 As shown, a neurosurgery navigation system 100 based on augmented reality technology according to an embodiment of the present application includes: a tracking data acquisition module 110, used to obtain marker point tracking data collected by an optical tracking system; an instrument data generation module 120, used to generate a real-time pose and type list of the tracked instrument in the current field of view based on the marker point tracking data; a surgery stage recognition module 130, used to input the real-time pose and type list of the tracked instrument in the current field of view into a surgery stage recognition model to obtain the currently estimated surgery stage; a target parameter data extraction module 140, used to extract a target visualization parameter set from a context-visualization rule database based on the currently estimated surgery stage; a visualization parameter prediction module 150, used to calculate a visualization parameter set used for rendering the next frame based on a comparison between the target visualization parameter set and the visualization parameter set at the current moment, and in combination with a smooth transition parameter; an AR navigation view generation module 160, used to render based on the visualization parameter set used for rendering the next frame to obtain an AR navigation view.

[0025] In the aforementioned augmented reality-based neurosurgery navigation system 100, the tracking data acquisition module 110 and the instrument data generation module 120 are used to acquire landmark tracking data collected by the optical tracking system and, based on this landmark tracking data, generate a real-time pose and type list of the tracked instruments within the current field of view. It should be understood that surgical instruments are the primary tools used by surgeons to perform procedures. Their position and orientation (i.e., pose) in three-dimensional space, as well as the instrument's type, directly reflect the surgeon's ongoing surgical actions and intentions. By accurately capturing this information in real time, the system can perceive the actual progress of the surgery, understand the current surgical stage, and provide highly relevant and timely navigation guidance. Furthermore, due to their non-contact, high-precision, and low-latency advantages, optical tracking systems are widely used in surgical navigation and are a common and reliable means of obtaining real-time instrument pose information. Fiducial markers, which serve as reference points that can be accurately identified and located by the optical tracking system, are fixed to surgical instruments, enabling the system to track the positions of these markers and calculate the overall pose of the attached instrument. At the same time, the specific landmark configurations or identification codes pre-arranged on different instruments also enable the system to distinguish different types of instruments. Therefore, obtaining these landmark data and parsing them into the real-time position and type list of instruments is a key step in connecting physical surgical operations with digital navigation systems.

[0026] More specifically, in one specific example of the present application, optical markers are first installed on the surgical instrument to be tracked in a predetermined configuration with a fixed geometric relationship. These markers can be reflective spheres or active or passive markers with specific patterns. Different marker configurations or markings can be used for different instrument types. Next, the camera array of the optical tracking system is positioned above or to the side of the surgical field, enabling it to clearly "see" the instruments with markers within the surgical field of view. During surgery, the cameras of the optical tracking system continuously capture images containing the markers at a high frame rate. After preprocessing the captured images, the system uses image processing algorithms to identify and accurately locate the positions of the markers in the images in a two-dimensional pixel coordinate system. Because the system includes multiple cameras with known relative positions and orientations (i.e., they have been calibrated), the positions of the markers from the two-dimensional pixel coordinates can be reconstructed into three-dimensional spatial coordinates in the optical tracking system coordinate system through multi-view geometry principles, such as triangulation. For each identified set of landmarks belonging to the same instrument, since their relative positions in the instrument's coordinate system are fixed (determined through preoperative or system initialization calibration), the system calculates the real-time 6-DOF pose (including 3D position and orientation) of the surgical instrument represented by these landmarks in the optical tracking system's coordinate system by solving a rigid-body transformation problem (e.g., using singular value decomposition or BundleAdjustment methods). Simultaneously, the system matches the identified landmark configurations or identities with a pre-defined database of instrument landmark configurations to determine the tracked instrument's type (e.g., aspirator, dissector, or bipolar electrocoagulation forceps). Finally, the system combines the real-time pose information (typically represented as a translation vector and a rotation matrix or quaternion) of all successfully tracked instruments at that moment, along with their identified types, into a list output: "The real-time pose and type list of tracked instruments in the current field of view," which is used by the recognition model in subsequent surgical phases.

[0027] In the aforementioned augmented reality-based neurosurgery navigation system 100, the surgical stage identification module 130 is configured to input the real-time position and type list of the tracked instruments within the current field of view into the surgical stage identification model to determine the current estimated surgical stage. It should be understood that existing neurosurgery navigation systems generally lack dynamic perception and understanding of the complex surgical process itself, making it difficult to intelligently adjust navigation information based on the actual progress and stage of the surgery. This results in information that is inconsistent with the surgeon's current needs and reduces navigation effectiveness. Surgical instruments, as the carriers of the surgeon's operations, their type, real-time position, and time-varying trajectory directly reflect the surgeon's ongoing operations and the stage of the surgery. Therefore, analyzing the real-time position and type list of instruments is a direct and feasible way to obtain dynamic information about the surgical process. Its purpose is clear and critical: to automatically and in real time identify the precise stage of the current surgery. Surgery is not a single, continuous action, but rather a series of stages with specific objectives and operational characteristics. Accurately understanding the current surgical stage is a prerequisite for intelligent and adaptive navigation. By inputting the collected real-time instrument pose and type data into the surgical stage recognition model, the system aims to extract high-level, surgically semantic stage information from these underlying, scattered physical data, thereby overcoming the shortcomings of traditional navigation systems that cannot understand the surgical process.

[0028] Figure 2 FIG. 1 is a block diagram of a surgical stage identification module in a neurosurgery navigation system based on augmented reality technology according to an embodiment of the present application. Figure 2 As shown, in an embodiment of the present application, the surgical stage identification module 130 includes: a data grouping unit 131, which is used to group the real-time posture and type list of the tracked instrument in the current field of view based on type to obtain a set of real-time posture time series of the tracking instrument; a posture timing pattern extraction unit 132, which is used to extract the device posture change timing pattern characteristics from each tracking instrument real-time posture time series in the set of tracking instrument real-time posture time series to obtain a set of tracking instrument posture change timing pattern feature coding vectors; a posture timing collaboration unit 133, which is used to perform posture timing collaborative aggregation analysis on the set of tracking instrument posture change timing pattern feature coding vectors to obtain a global dynamic collaborative semantic aggregation coding vector of the tracking instrument; a surgical stage estimation unit 134, which is used to determine the currently estimated surgical stage based on the global dynamic collaborative semantic aggregation coding vector of the tracking instrument.

[0029] In the aforementioned augmented reality-based neurosurgery navigation system 100, the data grouping unit 131 is configured to group the real-time poses and type lists of tracked instruments within the current field of view by type to obtain a collection of real-time pose time series of tracked instruments. It should be understood that the real-time pose and type list directly obtained from the optical tracking system is a snapshot of all tracked instruments within the current field of view. However, identifying surgical stages is not based on a single moment in time, but rather relies on the continuity and evolution of the surgeon's operation, particularly the motion patterns and coordinated behavior of different instruments over time. Different types of surgical instruments (such as aspirators, dissectors, and scissors) perform different functions in neurosurgery, and their typical usage and motion trajectories vary significantly. These differentiated behavior patterns are closely related to specific surgical stages. Simply mixing the raw instantaneous data of all instruments together would obscure the unique motion semantic information of these types. Therefore, the real-time poses and type lists of tracked instruments within the current field of view are further grouped by type to obtain a collection of real-time pose time series of tracked instruments. The mixed real-time data is structured and classified by instrument type. Through this grouping, the system can aggregate the posture data of instruments of the same type at consecutive time points to construct independent real-time posture time series divided by instrument type. This is done so that targeted temporal pattern feature extraction and analysis can be performed for each specific type of instrument, or to analyze the relationship between different instrument types. Only by grouping data by type can the unique motion characteristics exhibited by specific instrument types at different surgical stages, as well as the possible collaborative operation modes between different types of instruments, be effectively captured.

[0030] In the aforementioned augmented reality-based neurosurgery navigation system 100, the posture temporal pattern extraction unit 132 is configured to extract device posture change temporal pattern features from each tracking instrument real-time posture time series in the set of tracking instrument real-time posture time series to obtain a set of tracking instrument posture change temporal pattern feature encoding vectors. It should be understood that although the real-time instrument data has been grouped by type and organized into tracking instrument real-time posture time series, the raw tracking instrument real-time posture time series data remains low-level, high-dimensional, and may contain noise. The distinction between surgical stages is not solely based on the instantaneous position of the instrument or a single motion parameter, but rather relies more on the instrument's continuous motion pattern, manipulation technique, and trajectory characteristics over a period of time. For example, delicate peeling operations and rapid suction operations differ significantly in terms of instrument motion speed, acceleration, path complexity, and repeatability. Simply using the raw tracking instrument real-time posture time series data for subsequent analysis is inefficient and difficult to capture these complex temporal dynamic features. Therefore, in the technical solution of the present application, the device posture change timing pattern features are extracted from each tracking device real-time posture time series in the set of tracking device real-time posture time series to obtain a set of tracking device posture change timing pattern feature coding vectors.

[0031] In particular, in an embodiment of the present application, the real-time posture time series of each tracked instrument are processed separately through a device posture change temporal pattern feature extractor based on a time-series convolutional neural network to obtain a set of temporal pattern feature encoding vectors for the tracked instrument posture change. Processing by the temporal pattern feature extractor based on a time-series convolutional neural network automatically learns and extracts features that accurately characterize the instrument's motion pattern within the current time window. Notably, as a deep learning model, the time-series convolutional neural network is particularly adept at processing sequential data, effectively capturing local patterns, long-term dependencies, and complex correlations between different time steps in the time series. It can automatically learn multi-scale, multi-level temporal features from the real-time posture time series of the tracked instrument, such as the instrument's vibration frequency, motion smoothness, and specific trajectory patterns, and encode these features into a compact vector representation. This process converts the high-dimensional raw time series data into a low-dimensional but information-rich feature vector, namely the "tracked instrument posture change temporal pattern feature encoding vector," thereby effectively quantifying and characterizing the instrument's dynamic operational behavior and highlighting key dynamic features related to the type and intent of the surgical procedure.

[0032] In the aforementioned augmented reality-based neurosurgery navigation system 100, the posture temporal coordination unit 133 is configured to perform posture temporal coordination aggregation analysis on the set of tracking instrument posture change temporal pattern feature encoding vectors to obtain a global dynamic coordination semantics aggregation encoding vector for the tracking instruments. It should be understood that although the dynamic operation pattern features of each individual tracked instrument have been extracted and encoded from its real-time posture time series, resulting in a set of tracking instrument posture change temporal pattern feature encoding vectors reflecting the individual behavioral characteristics of each instrument, neurosurgery is often a complex multi-instrument collaborative operation process. Different surgical phases are not simply characterized by the motion of a single instrument, but more importantly, by the coordination, interaction, and overall dynamic patterns among multiple instruments. For example, an aspirator and a dissector may operate simultaneously, and their relative position changes and synchronous or asynchronous motion patterns together reveal the current stage of dissection and visual field clearing. Simply listing or combining these individual feature vectors fails to capture the complex high-order dependencies and overall coordination semantics between instruments. Therefore, in order to go beyond the individual feature analysis of a single instrument, in the technical solution of the present application, a posture temporal collaborative aggregation analysis is further performed on the set of tracking instrument posture change timing pattern feature coding vectors to obtain a global dynamic collaborative semantic aggregation coding vector for the tracking instrument. In this way, dynamic semantic information with globality and collaboration that reflects the behavior of the instrument group within the entire surgical field of view can be extracted and fused. This aggregation analysis aims to generate a single, high-information-density global dynamic collaborative semantic aggregation coding vector for tracking instruments. This vector not only contains information on the individual behavioral characteristics of each instrument, but more importantly, it can capture the collaborative features such as the interaction, synchronization, and relative motion pattern between them, thereby forming a more comprehensive and accurate overall description of the current surgical status. In this way, the system can extract global dynamic collaborative semantics that reflect the collaborative movement, mutual relationship, and overall operation mode of the instrument group from the tracking instrument posture change timing pattern feature coding vectors of scattered individual instruments. This generated "tracking instrument global dynamic collaborative semantic aggregate encoding vector" is a compact and information-dense representation that not only summarizes the average behavior or main trends of the entire instrument collection, but also highlights the unique operations of individual instruments or specific interaction patterns with other instruments that are important for distinguishing surgical stages. This capture of global collaborative semantics enables subsequent surgical stage recognition models to make judgments based on a more comprehensive and higher-level understanding of the scene, thereby more accurately identifying the current stage of the surgery. This overcomes the limitations of analyzing only individual instrument behavior and provides a solid semantic foundation for achieving true process adaptation and intelligent navigation.

[0033] Figure 3FIG is a block diagram of a posture timing coordination unit in a neurosurgery navigation system based on augmented reality technology according to an embodiment of the present application. Figure 3 As shown, in an embodiment of the present application, the posture timing coordination unit 133 includes: an instrument posture timing feature aggregation subunit 1331, which is used to perform a base state feature aggregation process based on cluster analysis on the set of tracking instrument posture change timing pattern feature coding vectors to obtain the tracking instrument posture change timing pattern set base state feature aggregation coding vector; a posture timing pattern excited state capture subunit 1332, which is used to calculate the position of each tracking instrument posture change timing pattern feature coding vector in the set of tracking instrument posture change timing pattern feature coding vectors relative to the tracking instrument posture change timing pattern based on the tracking instrument posture change timing pattern set base state feature aggregation coding vector. The excited state coding vector of the collective ground state feature aggregation coding vector is collected to obtain a set of tracking device posture change timing pattern feature compensation excited state coding vectors; the posture timing feature gain superposition subunit 1333 is used to determine the excited state significance modulation weight factor of each tracking device posture change timing pattern feature compensation excited state coding vector based on the set of tracking device posture change timing pattern feature compensation excited state coding vectors, and perform feature gain superposition on the set of tracking device posture change timing pattern feature compensation excited state significance coding vectors and the tracking device posture change timing pattern set ground state feature aggregation coding vector to obtain the tracking device global dynamic collaborative semantic aggregation coding vector.

[0034] Specifically, the device posture temporal feature aggregation subunit 1331 is used to perform cluster analysis-based ground state feature aggregation processing on the set of tracking device posture change temporal pattern feature coding vectors to obtain a set of ground state feature aggregation coding vectors of the tracking device posture change temporal pattern. Accordingly, in an embodiment of the present application, the device posture temporal feature aggregation subunit 1331 includes: a tracking device posture temporal clustering secondary subunit, used to input the set of tracking device posture change temporal pattern feature coding vectors into a K-Means clustering network to obtain K tracking device posture change temporal pattern initial cluster center coding vectors; a posture temporal set ground state feature aggregation secondary subunit, used to perform a self-attention mechanism-based set ground state feature aggregation on the K tracking device posture change temporal pattern initial cluster center coding vectors to obtain the tracking device posture change temporal pattern set ground state feature aggregation coding vector.

[0035] Specifically, the tracking device posture temporal clustering secondary subunit is used to input the set of tracking device posture change temporal pattern feature coding vectors into the K-Means clustering network to obtain K tracking device posture change temporal pattern initial cluster center coding vectors, which can be expressed as follows:

[0036]

[0037] in, is a set of characteristic coding vectors of the temporal pattern of the tracking device posture change, are respectively the first, second, and third in the set of the tracking device posture change time series pattern feature coding vectors. and A tracking device posture change time series pattern feature encoding vector, For K-Means clustering network processing, is the initial cluster center encoding vector of the K tracking device posture change time series pattern, They are respectively the first, second and kth initial cluster center coding vectors of the K tracking device posture change time series pattern initial cluster center coding vectors.

[0038] It should be understood that while each of the tracking device's temporal pattern encoding vectors encodes the dynamic behavior of a single device, this set may contain a large number of high-dimensional features, making direct, complex, global collaborative analysis computationally intensive and inefficient. Furthermore, these feature vectors may contain some redundancy, or it may be necessary to identify representative data pattern primitives that summarize different operating modes. Before conducting more refined collaborative aggregation analysis, these individual feature vectors require preliminary structured summarization and information compression to identify the primary, representative clusters of motion patterns within the set. Therefore, in this solution, the classic K-Means clustering algorithm is used to perform preliminary unsupervised learning and structured exploration on the input set of tracking device temporal pattern encoding vectors. K-Means aims to cluster similar feature vectors and calculates the centroid of each cluster to represent the average characteristics of all samples within that cluster. Through this process, the system can identify K primary operating modes or device behavior primitives within the set and represent the centers of these patterns as "K initial cluster center encoding vectors of tracking device temporal pattern encodings." These cluster centers can be viewed as a preliminary summary and compressed representation of the entire device behavior feature set, capturing the key, high-density feature information within the set. Its core function is data condensation and structured detection, compressing the massive original feature vector set into K representative pattern primitives.

[0039] Specifically, the posture temporal set ground state feature aggregation secondary subunit is used to perform a self-attention-based set ground state feature aggregation on the K initial cluster center encoding vectors of the tracking instrument posture change temporal pattern to obtain the set ground state feature aggregation encoding vector of the tracking instrument posture change temporal pattern. It should be understood that although K-Means clustering has identified K representative pattern primitives (the initial cluster centers of the tracking instrument posture change temporal pattern), these centers are merely the centroids of local data density regions. They exist independently and do not explicitly model the relationships between these pattern primitives or how they collectively constitute the overall behavioral pattern of the entire instrument set. Instrument operations in surgical scenarios are often complex, and different operation pattern primitives may have high-order dependencies or structural associations. For example, certain operation modes often occur simultaneously or alternately with other modes. Simply combining these independent cluster centers cannot capture this global, inter-pattern synergy information, nor can it easily extract "ground state" features that represent the commonalities of the entire set. Therefore, the technical solution of this application utilizes an aggregation method based on the self-attention mechanism to deeply explore and integrate the intrinsic connections and global structural information between these K initial cluster center encoding vectors. The self-attention mechanism allows each element in the network (here, each cluster center encoding vector) to pay attention to all other elements in the set and assign different weights based on their interrelationships. Through this mechanism, the network is able to model high-order dependencies between these cluster centers, for example, identifying spatially separate but semantically related operational mode primitives and understanding how they influence and complement each other to collectively constitute the current surgical state. Ultimately, by aggregating these contextually enhanced and globally informed cluster center representations, a single vector is extracted that reflects the intrinsic structure, core distribution characteristics, and collective common behavior of the entire instrument set. This vector is the "aggregated encoding vector of the ground state features of the temporal pattern of the tracked instrument posture changes." This ground state feature encoding vector is intended to serve as a representative summary of the dynamic behavior of the entire instrument set, capturing its "collective or global characteristics" and improving understanding of its collaborative behavior.

[0040] More specifically, the posture timing set ground state feature aggregation secondary sub-unit includes: performing self-attention encoding on each tracking device posture change timing pattern initial clustering center encoding vector of the K tracking device posture change timing pattern initial clustering center encoding vectors to obtain K tracking device posture change timing pattern self-attention encoding vectors; performing position-based aggregation on the K tracking device posture change timing pattern self-attention encoding vectors and the K tracking device posture change timing pattern initial clustering center encoding vectors to obtain the tracking device posture change timing pattern set ground state feature aggregation encoding vector.

[0041] Specifically, self-attention encoding is performed on each of the K tracking device posture change timing pattern initial cluster center encoding vectors to obtain K tracking device posture change timing pattern self-attention encoding vectors, which can be expressed as follows:

[0042]

[0043]

[0044]

[0045] in, 、 and They are respectively the tracking device posture change timing mode query weight matrix, the tracking device posture change timing mode key weight matrix and the tracking device posture change timing mode value weight matrix. 、 and are respectively the set of tracking device posture change timing pattern query vectors, the set of tracking device posture change timing pattern key vectors and the set of tracking device posture change timing pattern value vectors. for The scale, for function, is the self-attention mechanism, A collection of self-attention encoding vectors for tracking the temporal pattern of device posture changes.

[0046] More specifically, the K tracking device posture change time series pattern self-attention encoding vectors and the K tracking device posture change time series pattern initial cluster center encoding vectors are aggregated by position to obtain the tracking device posture change time series pattern set ground state feature aggregation encoding vector, which is expressed as:

[0047] in, To calculate and Positional addition of is the layer normalization process, Aggregate encoding vectors of ground state features for tracking device posture change temporal patterns.

[0048] Specifically, the posture timing pattern excited state capture subunit 1332 is used to calculate the excited state coding vector of each tracking device posture change timing pattern feature coding vector in the set of tracking device posture change timing pattern feature coding vectors relative to the ground state feature aggregate coding vector of the tracking device posture change timing pattern set based on the ground state feature aggregate coding vector of the tracking device posture change timing pattern set to obtain a set of tracking device posture change timing pattern feature compensation excited state coding vectors, which can be expressed as follows:

[0049]

[0050] in, and are the modulated trainable weight matrix and the modulated trainable bias vector, is the point product by position, for function, To track the temporal pattern characteristics of the device posture change, a set of excited state coding vectors is compensated. are the first, second, and third vectors in the set of the tracking device posture change timing pattern characteristic compensation excited state coding vectors. and The tracking device posture change timing pattern feature compensation excitation state encoding vector.

[0051] It should be understood that although the aggregated encoding vector of the collective ground state features representing the collective commonality or global characteristics of the entire instrument set is obtained through aggregation based on the self-attention mechanism, this only captures the average behavior or main trends of the set. In complex surgical procedures, fully understanding the deep semantics of posture-temporal coordination requires not only knowing the overall "ground state" but also understanding how the dynamic behavior of each individual instrument deviates from or contributes to this overall ground state. For example, in a coordinated operation, a single instrument may perform a key auxiliary action that slightly deviates from the overall rhythm. This individual deviation is crucial for distinguishing subtle surgical sub-phases. Relying solely on the ground state vector fails to capture this critical individual-specific information, which is an integral part of complex collaborative behavior. Therefore, to explicitly separate and quantify the portion of the feature encoding vector of each raw tracked instrument's posture change temporal pattern that is not fully explained by the collective ground state features, this differential information is extracted in the form of an "excitation state encoding vector." Its essence lies in isolating the deviation of each tracking device's posture change time series pattern feature encoding vector from the collective average behavior, and interpreting it as the "excitation state" or "perturbation" of the original tracking device posture change time series pattern feature. By calculating the difference or relative relationship between each individual tracking device posture change time series pattern feature encoding vector and the ground state feature aggregation encoding vector of the tracking device posture change time series pattern representing the collective commonality, the aim is to clearly capture and quantify the unique contribution or deviation of each device in the current collaborative operation, thereby obtaining a collection representing individual-specific information, thereby providing high-quality excited state information for subsequent aggregation.

[0052] Specifically, the posture timing feature gain superposition subunit 1333 is used to determine the excited state significance modulation weight factor of each tracking device posture change timing pattern feature compensation excited state coding vector based on the set of tracking device posture change timing pattern feature compensation excited state coding vectors, and perform feature gain superposition on the set of tracking device posture change timing pattern feature compensation excited state significance coding vectors and the tracking device posture change timing pattern set base state feature aggregation coding vector to obtain the tracking device global dynamic collaborative semantic aggregation coding vector. More specifically, the posture timing feature gain superposition subunit includes: calculating the excited state significance modulation weight factor of each tracking device posture change timing pattern feature compensation excited state coding vector in the set of tracking device posture change timing pattern feature compensation excited state coding vectors to obtain a set of tracking device posture change timing pattern exciting state significance modulation weight factors; based on the set of tracking device posture change timing pattern exciting state significance modulation weight factors, weighted modulating the set of tracking device posture change timing pattern feature compensation excited state coding vectors to obtain a set of tracking device posture change timing pattern feature compensation excited state significant coding vectors; inputting the set of tracking device posture change timing pattern feature compensation excited state significant coding vectors and the tracking device posture change timing pattern set base state feature aggregation coding vector into the feature gain superposition network to obtain the tracking device global dynamic collaborative semantic aggregation coding vector.

[0053] Specifically, the excited-state significance modulation weight factor of each tracking device posture change timing pattern feature compensation excited-state coding vector in the set of tracking device posture change timing pattern feature compensation excited-state coding vectors is calculated to obtain a set of tracking device posture change timing pattern excited-state significance modulation weight factors, which is expressed as follows:

[0054]

[0055] in, To stimulate the modulation of the trainable weight matrix, for activation function, To stimulate the modulation trainable vector, is the exponential function value with the natural constant e as the base, To track the set of significant modulation weight factors of the excited state of the time series pattern of the device posture change, are respectively the first, second, and third in the set of significant modulation weight factors of the excited state of the timing pattern of the tracking device posture change. and The significance modulation weight factor of the excited state of the temporal pattern of the posture change of a tracking device.

[0056] It should be understood that not all individual deviations in the set of excitation state encoding vectors for tracking instrument posture changes have equal surgical semantic importance. In actual surgical procedures, some instrument excitation states may represent random noise, minor jitter, or incidental movements unrelated to the current surgical stage, which should not be given excessive attention. However, other instrument excitation states may represent critical, intentional auxiliary actions or fine-tuning performed by the surgeon. These deviations are crucial for accurately identifying surgical sub-stages. If these excitation states are not differentiated and all individual deviations are lumped together and directly used for subsequent aggregation, noise may be introduced, diluting critical information and affecting the accuracy of the final surgical stage identification. Therefore, a mechanism is needed to evaluate and quantify the significance or importance of each excitation state. Therefore, the set of excitation state encoding vectors for tracking instrument posture changes, which are used to compensate for the temporal pattern of excitation states, is further salience modulated. By calculating the excitation state saliency modulation weight factors corresponding to the excitation state encoding vectors compensated for the temporal pattern characteristics of each tracked device's posture changes, the system aims to distinguish meaningful individual deviations (i.e., excitation states with high saliency) from random noise or unimportant details (i.e., excitation states with low saliency). These weight factors reflect the potential contribution or information content of each individual excitation state to the overall collaborative semantics. By calculating these weights, the system can quantify the importance of the individual specificity exhibited by each device in the current operation, preparing for subsequent information fusion and ensuring that only truly discriminative excitation state information that reflects key individual operations can participate in the final aggregation with higher weights.

[0057] Specifically, based on the set of significant modulation weight factors of the excited state of the tracking device posture change timing pattern, the set of the tracking device posture change timing pattern feature compensation excited state coding vectors is weighted modulated to obtain the set of significant coding vectors of the tracking device posture change timing pattern feature compensation excited state, which can be expressed as follows:

[0058]

[0059] in, To track the temporal pattern characteristics of the device posture change, the set of excited state significant coding vectors is compensated. are respectively the first, second, and third significant coding vectors in the set of the tracking device posture change time series pattern feature compensation excited state. and The temporal pattern characteristics of the tracking device's posture change compensate the excited state saliency encoding vector.

[0060] It should be understood that although the previous step has calculated the excitation-state encoding vectors for the temporal pattern of tracking instrument posture changes relative to the group base state, these vectors represent the deviation of each instrument's motion pattern from the overall trend. However, not all individual deviations have equal surgical semantic importance; some may be random noise, minor operator jitter, or background motion unrelated to the current critical collaborative action. When analyzing multi-instrument "posture temporal collaboration," the goal is to capture those individual differences that are critical for distinguishing different collaborative patterns or surgical stages, rather than all minor fluctuations. Simply treating all excitation-state encoding vectors equally and directly using them for subsequent aggregation will introduce noise, dilute the truly critical individual contribution information, and thus affect the accurate understanding of complex collaborative patterns. Therefore, in the technical solution of this application, these excitation-state encoding vectors for tracking instrument posture change temporal pattern feature compensation are further processed with discriminativeness using the excitation-state saliency modulation weight factor for tracking instrument posture change temporal pattern feature compensation. By multiplying each tracking instrument posture change temporal pattern feature compensation excitation-state encoding vector by its corresponding saliency weight factor, the goal is to "weighted modulate" the individual deviation information. Excitation states assessed as having high significance (i.e., deviations that are considered to be more representative of meaningful individual operations or significant contributions to collaborative behavior) receive higher weights, thereby emphasizing and retaining their characteristic information after weighting; whereas excitation states assessed as having low significance (possibly representing noise or unimportant details) receive lower weights, thereby suppressing or attenuating their information after weighting. This salience modulation of individual excitation states ensures that the final global dynamic collaborative semantic aggregate encoding vector can focus more on capturing the most discriminative collaborative patterns and key individual contributions within the instrument population, significantly improving the depth and accuracy of understanding the temporal collaborative patterns of complex surgical postures, and providing purer, more informative, and more targeted input for subsequent surgical stage identification.

[0061] Specifically, the set of the tracking device posture change temporal pattern feature compensation excitation state significant coding vectors and the tracking device posture change temporal pattern set ground state feature aggregate coding vector are input into the feature gain superposition network to obtain the tracking device global dynamic collaborative semantic aggregate coding vector, which is expressed as follows:

[0062] in, The number of vectors in the set of significant coding vectors for compensating the excited state of the tracking device posture change timing pattern feature, is a multi-layer perceptron, A global dynamic collaborative semantic aggregation encoding vector is generated for the tracking device.

[0063] It should be understood that although the previous steps have generated a ground state feature aggregation encoding vector representing the collective commonality or ground state of the entire instrument ensemble, and a set of significant encoding vectors for compensated excited state features of the temporal pattern of tracked instrument pose changes relative to the ground state, which quantifies and selects significant individual instrument pose changes relative to the ground state, these two types of information (global commonality and significant individual differences) remain separate representations. To fully understand the deep semantics of instrument group pose temporal coordination and integrate them into a single, compact representation applicable to downstream tasks (such as surgical stage recognition), we cannot rely solely on a simple listing of ground state information or excited state sets. The complexity of surgery lies in the individual key behaviors within the overall context and how multiple individual behaviors organically combine to form specific collaborative patterns. A mechanism is needed to organically integrate the ground state information representing overall trends with the excited state information representing the deviations and unique contributions of key individuals, forming a comprehensive representation that combines a global perspective while retaining local details. By superimposing the aggregated feature gain of the ground state feature vector representing the common temporal patterns of tracked device pose changes with the salient feature vectors representing the compensated excitation states of the temporal patterns of tracked device pose changes (which have been screened and weighted), the network aims to construct a final, information-complete aggregate representation. By organically integrating the ground state vectors representing commonalities with the set of excitation state vectors representing significant individual differences, a single aggregated encoding vector for the global dynamic collaborative semantics of the tracked devices is generated. This vector is designed to compactly encapsulate the multi-level information of the entire input set: it encompasses both the overall average profile or main trends (ground state reflection) and highlights the most noteworthy individual highlights or key auxiliary behaviors (significant excitation state reflection). Specifically, it integrates the ground state information representing the overall dynamic trends of the device group with the excitation state information representing the key operations or interactions of individual devices with surgical semantic significance. This aggregated encoding vector provides a comprehensive, abstract, and discriminative representation of the collaborative behavior patterns of all tracked devices within the current surgical field of view, capturing the complex collaborative semantics that are difficult to express with individual device features or simple combinations. For example, it can distinguish whether the surgeon is simply picking up two instruments or using them for a coordinated peeling and suctioning operation. This high-quality global dynamic collaborative semantic aggregate encoding vector of tracked instruments is fed into a classifier-based surgical stage recognition model as the final input. This enables the model to more accurately and robustly identify the precise stage of the current surgery based on a deep understanding of the collaborative dynamics of the instrument group, greatly enhancing the intelligent perception and navigation capabilities of the entire system.

[0064] Preferably, when calculating the activation state representation of the characteristic coding vector of each tracking device posture change time series pattern relative to the base state characteristic aggregation coding vector of the tracking device posture change time series pattern set, Essentially defines the ground state space induced metric of each tracking device posture change time series pattern feature coding vector relative to the tracking device posture change time series pattern set ground state feature aggregation coding vector, that is, the relatively high-dimensional original feature vectors The feature space of the tracking device is used as a measurement benchmark to define a low-dimensional space measurement standard of the ground state feature aggregation encoding vector of the set of temporal pattern sets of the tracking device posture change through guided mapping.

[0065] In this way, it is necessary to consider the boundary measurement problem under the induced metric standard, that is, to make the measurement of the activation state subject to the geometric constraints of the spatial transition under the guided mapping, rather than a simple high-dimensional-low-dimensional space transformation.

[0066] Based on this, first calculate Relative to The low-rank activation coefficient and , and introduce the low-rank tangent space form and , thereby solving the guided mapping in the form of tangent space so that it can be projected onto an orthogonal basis in the basis state space.

[0067] Then, the low-rank tangent space form is used as a metric representation under the local submanifold structure to reconstruct the boundary-space geometry as follows:

[0068] The metric complexity of the excited state is transformed from being determined by the high-dimensional set representation of the temporal pattern feature encoding vector of the tracking device posture change to being determined by the boundary geometric characteristics of the tangent space metric standard. Is the adjustment coefficient, used to compensate for excessive spatial geometric parameters The impact brought about.

[0069] In this way, Correction for , under the induced metric standard of the submanifold metric standard to the ground state, by making the activation metric depend on the boundary rather than the spatial transformation, the normalization of the fixed points in the tangent direction metric flow can be realized to achieve effective measurement of the boundary and avoid excessive expansion of the representation of the excited state encoding vector compensated by the temporal pattern characteristics of the posture change of each tracking device.

[0070] In the aforementioned augmented reality-based neurosurgery navigation system 100, the surgical stage estimation unit 134 is configured to determine the currently estimated surgical stage based on the global dynamic collaborative semantic aggregate coding vector of the tracked instruments. It should be understood that the global dynamic collaborative semantic aggregate coding vector of the tracked instruments encodes complex collaborative semantic information that organically integrates the global commonalities and key individual differences of the instrument group dynamics within the current time window with high information density. This vector represents a comprehensive, multi-level understanding of the dynamic behavior of the instrument group. The distinction between surgical stages is largely based on the specific, time-evolving collaborative operation patterns of these instrument groups. Therefore, the global dynamic collaborative semantic aggregate coding vector of the tracked instruments is the input signal that currently best represents the core characteristics of the surgical stage and is the most direct and effective basis for stage determination. Using raw data or simple feature combinations cannot achieve the same effect. Therefore, in an embodiment of the present application, the global dynamic collaborative semantic aggregate coding vector of the tracked instruments is input into a classifier-based surgical stage recognition model to obtain the currently estimated surgical stage. The classifier-based surgical stage recognition model's responsibility is to perform pattern recognition and classification on the input global dynamic collaborative semantic aggregate coding vector of the tracked instruments based on a pre-trained model. This classifier model essentially learns the distribution or mapping of the global dynamic collaborative semantic aggregate encoding vectors of tracked instruments corresponding to different surgical stages. Its purpose is to map the input global dynamic collaborative semantic aggregate encoding vector of the tracked instrument to a discrete, predefined surgical stage label (e.g., "exposure," "dissection," "suture," etc.), thereby outputting the system's estimate of the current surgical stage. This step transforms the complex, continuous instrument dynamic behavior signal into a clear stage identifier with semantically relevant surgical process information.

[0071] In the aforementioned augmented reality-based neurosurgery navigation system 100, the target parameter data extraction module 140 is configured to extract a set of target visualization parameters from the context-visualization rule database based on the currently estimated surgical stage. It should be understood that neurosurgery is a dynamically evolving process, with different surgical stages corresponding to completely different operational tasks, critical anatomical structures, and potential risk points. For example, during the craniotomy phase, the focus may be on bone boundaries and vascular distribution; after dural opening, the focus shifts to cortical anatomical structures and tumor boundaries; and during the hemostasis phase, precise identification of bleeding points and associated blood vessels is required. The previous step successfully maps the dynamic collaborative behavior of real-time instruments to an accurate surgical stage estimate through complex feature extraction and aggregation analysis, providing the system with a high-level semantic understanding of the current surgical process. For augmented reality (AR) navigation systems to truly and effectively assist surgeons, the information they overlay during surgery must be closely relevant to the surgeon's current task, non-intrusive, and highly valuable. Static, unchanging AR displays fail to meet this requirement and may even distract the surgeon by displaying too much irrelevant information. Therefore, dynamically adjusting the AR visualization content and style based on the contextual information of the current surgical stage is key to achieving intelligent, process-adaptive navigation. Notably, this parameter set specifies in detail how the AR display should be presented during the current surgical stage. This includes, but is not limited to, which 3D anatomical structures (such as tumors, blood vessels, nerves, and bones) should be displayed or hidden, how these structures should be rendered (transparency, color, wireframe, solid), how key areas or targets should be highlighted (flashing, bold outlines, color changes, etc.), whether the surgical plan trajectory or safety margins should be displayed, whether specific text or warnings should be overlaid, and other visualization-related configuration options. By extracting these stage-specific visualization parameters, the system ensures that the AR display accurately reflects the current surgical task and the surgeon's focus, providing the most relevant and intuitive navigation information.

[0072] Specifically, in one specific example of this application, the steps for extracting a target visualization parameter set from a context-visualization rule database are as follows: the system first receives a discrete label representing the current estimated surgical stage (e.g., a string or enumeration value, such as "dura mater exposure stage" or "tumor boundary identification stage") output by a classifier-based surgical stage recognition model. The system then accesses its internal or externally stored "context-visualization rule database." This database is a structured storage containing multiple records, each of which maps a specific surgical stage label to a corresponding set of visualization parameter configurations. The system uses the received surgical stage label as a query key and performs a search operation within the database to locate a record that exactly matches the label. Once a matching record is found, the system extracts a predefined set of visualization parameters from that record. These parameters are typically stored in a structured format, such as a JSON object, XML file, or instance of a specific configuration class. They contain all the information required by the AR rendering engine to guide how to draw the 3D model, overlay information, set styles, and so on. The extracted "target visualization parameter set" is then passed to the system's AR rendering module or visualization manager. The AR rendering module dynamically configures and updates its rendering pipeline based on the received visualization parameter set, and adjusts the content, style, and interactive behavior of the AR overlay display in real time to accurately match the currently estimated surgical stage.

[0073] In the aforementioned augmented reality-based neurosurgery navigation system 100, the visualization parameter prediction module 150 is configured to calculate the visualization parameter set used for rendering the next frame based on a comparison between the target visualization parameter set and the current visualization parameter set, combined with a smooth transition parameter. It should be understood that the recognition results of surgical stages may fluctuate slightly. Even if the recognition results are stable, if the AR display content changes abruptly when switching from one surgical stage to the next, such as certain structures suddenly appearing or disappearing, or attributes such as color or transparency changing instantly, this can cause visual disturbance to the surgeon, distracting their attention and potentially even causing confusion. A high-quality AR navigation system should provide a smooth and natural visual experience, so that even when switching stages or when the recognition results fluctuate slightly, the AR content updates in a gradual manner to avoid abruptness. Therefore, a mechanism is needed to manage the changes in AR visualization parameters over time to ensure a smooth and controllable transition from the current state to the target state. Therefore, in order to achieve a smooth transition of AR visualization parameters, by comparing the visualization parameter set used for AR rendering at the current moment (i.e., the "visualization parameter set at the current moment") with the "target visualization parameter set" extracted based on the latest surgical stage, the system can determine which parameters need to be changed and in what direction. Then, combined with the preset "smooth transition parameters" (for example, the duration of the transition, the type of easing curve, etc.), the system calculates the visualization parameter set of the intermediate state for rendering the next frame. This intermediate parameter set is an interpolation result between the current parameter set and the target parameter set, which allows the AR display to gradually approach the target state in time rather than jumping instantly. The purpose of this is to ensure that updates to the AR interface can reflect the latest surgical stage information without disturbing the surgeon due to abrupt changes, thereby improving the system's usability and user experience.

[0074] Specifically, in a specific example of the present application, the "current visualization parameter set" currently being used for AR rendering is first obtained. This parameter set records all visualization configurations used when rendering the previous frame, such as the current transparency, color value, and visibility status of an anatomical structure. The system also obtains the "target visualization parameter set" extracted from the database in the previous step based on the most recently identified surgical stage. This parameter set represents the final state of the AR display corresponding to the current surgical stage under ideal circumstances. The system compares the "current visualization parameter set" with the "target visualization parameter set" and identifies all parameter items that are different. For example, if the goal is to change a structure from transparent to opaque, or from invisible to visible. Further, for each parameter item that is different, the system will start a transition process. The system maintains the current transition status of each parameter item, including the transition start time, starting value, target value, and remaining transition time. The system then retrieves the preset "smooth transition parameters," which typically include a global transition duration (e.g., a parameter transitions in 0.5 seconds) or different durations for specific parameter types, as well as possible easing functions (e.g., linear, exponential, sine, etc.). When calculating the parameters for the next frame, the system considers the time elapsed since the transition began. Based on the elapsed time, the total transition duration, and the selected easing function, the system calculates an interpolation factor (typically between 0 and 1). This interpolation factor represents the progress of the current transition. For each parameter being transitioned, the system uses its starting value, target value, and the calculated interpolation factor to calculate the parameter's new value for the next frame using an interpolation algorithm (e.g., linear interpolation: current value = starting value * (1 - interpolation factor) + target value * interpolation factor). For Boolean visibility parameters, more complex logic may be required, such as toggling visibility when the interpolation factor reaches a threshold (e.g., 0.5) or adjusting transparency to achieve a fade-in / fade effect. All new parameter values ​​obtained through interpolation calculation are collected to form the "visualization parameter set used for the next frame rendering" for the next frame of AR rendering. For those parameters that do not differ from the target parameter set at the current moment, their values ​​remain unchanged. This calculated parameter set is then passed to the AR rendering engine for rendering the next frame of AR image. At the same time, the "visualization parameter set used for the next frame rendering" will be updated to the new "visualization parameter set at the current moment" to prepare for the calculation of the next frame.

[0075] In the aforementioned augmented reality-based neurosurgery navigation system 100, the AR navigation view generation module 160 is configured to render an AR navigation view based on the visualization parameter set used for the next frame rendering. It should be understood that the visualization parameter set used for the next frame rendering is simply a set of numerical configuration information that describes which virtual objects should be displayed, where, and in what format at the current moment. This abstract parameter information must be transformed into an image through a visualization process before it can be truly presented to the surgeon and fulfill its navigational function. Rendering is the key process in transforming numerical configurations into visual images. Through the graphics rendering engine, information such as the virtual 3D anatomical model, surgical planning path, safety margins, and text prompts are precisely rendered and overlaid onto the real-time surgical field of view. This process aims to generate a final augmented reality view. By observing this view, the surgeon can simultaneously see the real anatomical structure and the precisely aligned and highly relevant virtual navigation information, thereby assisting in surgical decision-making and operation. The goal is to "materialize" complex computational results into intuitive visual cues, bridging the gap between abstract data and actual surgical operations.

[0076] Specifically, in one specific example of this application, the system first obtains a real-time surgical field image stream captured by an AR device (such as a head-mounted display with a camera or an AR surgical microscope). Simultaneously, the system uses a tracking system (e.g., an optical tracker) to obtain the precise 3D pose (position and orientation) of the AR device's camera relative to the patient's anatomy or a predefined surgical space coordinate system. This pose information is critical for achieving precise alignment of virtual content with the real world. The system then receives the "visualization parameter set for the next frame rendering" calculated in the previous step. This parameter set contains all the configuration information required for rendering the current frame, such as the ID of the anatomical models to be loaded and displayed, their spatial transformations (typically based on preoperative image registration results), material properties (color, transparency, texture), visible landmarks, highlight areas, text overlay content, line thickness, flicker frequency, and so on. All of this is determined based on the current estimated surgical phase and transition state. Based on the received camera pose, the graphics rendering engine sets the viewing angle and projection parameters of the virtual camera in the virtual 3D scene, simulating how the surgeon observes the real world. The rendering engine loads and positions the virtual 3D models (such as blood vessels, tumors, and important nerve bundles) and surgical planning elements (such as approach trajectories and target areas) to be displayed into the virtual scene. Their spatial position and orientation are typically based on preoperative planning and intraoperative image registration. The rendering engine applies corresponding visualization properties to these virtual 3D models and elements based on the visualization parameter set used for the next frame rendering. For example, the rendering engine can set the transparency and color of specific models based on these parameters; enable or disable the visibility of certain models; highlight specific areas; overlay text labels at designated locations; and set the rendering style (such as wireframe or solid). The rendering engine performs rendering calculations, generating a virtual image frame from the perspective of the virtual camera. This image frame contains only virtual navigation elements, but their position and scale accurately correspond to the real world. Finally, the system composites (overlays) this virtual image with the live surgical field image captured at the current moment. This compositing process typically involves color blending (for example, transparent overlay based on the alpha channel) to ensure that the virtual content blends naturally with the real background, generating the final AR navigation view. The generated AR navigation view is then displayed on the display screen of the AR device for the doctor to observe and use in real time, thereby providing intelligent navigation assistance based on dynamic collaborative analysis of instruments and identification of surgical stages.

[0077] In summary, according to the embodiment of the present application, a neurosurgery navigation system based on augmented reality technology is explained, which can utilize the real-time dynamic perception and understanding of surgical instruments to automatically identify the current surgical stage, and use this stage information as a driving force to search in a context-visualization rule database that has been pre-built and stores visualization rules corresponding to different surgical stages, and extract a complete set of target visualization parameter sets that best matches the specific stage, thereby dynamically adjusting the content and presentation of the AR navigation view.

[0078] Furthermore, a neurosurgery navigation method based on augmented reality technology is also provided.

[0079] Figure 4 Flowchart of a neurosurgery navigation method based on augmented reality technology according to an embodiment of the present application. Figure 5 Schematic diagram of data flow of the neurosurgery navigation method based on augmented reality technology according to an embodiment of the present application. Figure 4 and Figure 5 As shown, the neurosurgery navigation method based on augmented reality technology according to an embodiment of the present application includes: S100, acquiring landmark tracking data collected by an optical tracking system; S200, generating a real-time pose and type list of the tracked instrument in the current field of view based on the landmark tracking data; S300, inputting the real-time pose and type list of the tracked instrument in the current field of view into a surgical stage recognition model to obtain the currently estimated surgical stage; S400, extracting a target visualization parameter set from a context-visualization rule database based on the currently estimated surgical stage; S500, calculating a visualization parameter set used for rendering the next frame based on a comparison between the target visualization parameter set and the visualization parameter set at the current moment and in combination with a smooth transition parameter; S600, rendering based on the visualization parameter set used for rendering the next frame to obtain an AR navigation view.

[0080] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A neurosurgery navigation system based on augmented reality technology, characterized in that: include: Tracking data acquisition module, used to obtain the marker tracking data collected by the optical tracking system; An instrument data generation module, configured to generate a real-time position and type list of the tracked instrument within the current field of view based on the landmark tracking data; a surgical stage recognition module, configured to input the real-time position and type list of the tracked instrument in the current field of view into a surgical stage recognition model to obtain a currently estimated surgical stage; a target parameter data extraction module, configured to extract a target visualization parameter set from a context-visualization rule database based on the currently estimated surgical stage; A visualization parameter prediction module, configured to calculate a visualization parameter set used for rendering the next frame based on a comparison between the target visualization parameter set and the visualization parameter set at the current moment and in combination with a smooth transition parameter; An AR navigation view generation module is configured to perform rendering based on the visualization parameter set used for rendering the next frame to obtain an AR navigation view.

2. The neurosurgery navigation system based on augmented reality technology according to claim 1, characterized in that: The surgical stage identification module includes: a data grouping unit, configured to group the real-time position and type list of the tracked device within the current field of view based on type to obtain a set of time series of the real-time position and type of the tracked device; a posture time series pattern extraction unit, configured to extract a device posture change time series pattern feature from each tracking device real-time posture time series in the set of tracking device real-time posture time series to obtain a set of tracking device posture change time series pattern feature encoding vectors; A posture time sequence coordination unit is used to perform posture time sequence coordination aggregation analysis on a set of characteristic coding vectors of the tracking device posture change time sequence pattern to obtain a global dynamic coordination semantic aggregation coding vector of the tracking device; The surgical stage estimation unit is used to determine the currently estimated surgical stage based on the global dynamic collaborative semantic aggregation coding vector of the tracking instrument.

3. The neurosurgery navigation system based on augmented reality technology according to claim 2, characterized in that: The posture timing pattern extraction unit is used to: pass the real-time posture time series of each tracking device through a device posture change timing pattern feature extractor based on a timing convolutional neural network to obtain a set of tracking device posture change timing pattern feature coding vectors.

4. The neurosurgery navigation method based on augmented reality technology according to claim 3, characterized in that: The posture timing coordination unit includes: An instrument posture temporal feature aggregation subunit, configured to perform a base state feature aggregation process based on cluster analysis on the set of the tracking instrument posture change temporal pattern feature coding vectors to obtain a tracking instrument posture change temporal pattern set base state feature aggregation coding vector; a posture timing pattern excited state capturing subunit, configured to calculate, based on the base state feature aggregation coding vector of the tracking device posture change timing pattern set, the excited state coding vector of each tracking device posture change timing pattern feature coding vector in the set of the tracking device posture change timing pattern feature coding vector relative to the base state feature aggregation coding vector of the tracking device posture change timing pattern set, so as to obtain a set of tracking device posture change timing pattern feature compensation excited state coding vectors; The posture timing feature gain superposition subunit is used to determine the excited state significance modulation weight factor of each tracking device posture change timing pattern feature compensation excited state coding vector based on the set of tracking device posture change timing pattern feature compensation excited state coding vectors, and perform feature gain superposition on the set of tracking device posture change timing pattern feature compensation excited state significance coding vectors and the tracking device posture change timing pattern set base state feature aggregation coding vector to obtain the tracking device global dynamic collaborative semantic aggregation coding vector.

5. The neurosurgery navigation system based on augmented reality technology according to claim 4, characterized in that: The device posture temporal feature aggregation subunit includes: The tracking device posture temporal clustering secondary subunit is used to input the set of tracking device posture change temporal pattern feature coding vectors into the K-Means clustering network to obtain K tracking device posture change temporal pattern initial clustering center coding vectors; The second-level sub-unit for posture temporal set basis state feature aggregation is used to perform collective basis state feature aggregation based on the self-attention mechanism on the initial cluster center encoding vectors of the K tracking device posture change temporal pattern to obtain the collective basis state feature aggregation encoding vector of the tracking device posture change temporal pattern.

6. The neurosurgery navigation system based on augmented reality technology according to claim 5, characterized in that: The posture time series aggregate ground state feature aggregation secondary subunit is used to: Performing self-attention encoding on each of the K tracking device posture change time series pattern initial cluster center encoding vectors to obtain K tracking device posture change time series pattern self-attention encoding vectors; The K tracking device posture change time series pattern self-attention coding vectors and the K tracking device posture change time series pattern initial cluster center coding vectors are aggregated by position to obtain the tracking device posture change time series pattern set ground state feature aggregation coding vector.

7. The neurosurgery navigation system based on augmented reality technology according to claim 6, characterized in that: The posture temporal feature gain superposition subunit is used to: Calculating the excited-state significance modulation weight factor of each tracking device posture change timing pattern feature compensation excited-state coding vector in the set of tracking device posture change timing pattern feature compensation excited-state coding vectors to obtain a set of tracking device posture change timing pattern excited-state significance modulation weight factors; Based on the set of significant modulation weight factors of the excited state of the tracking device posture change timing pattern, weighted modulation is performed on the set of characteristic compensation excited state coding vectors of the tracking device posture change timing pattern to obtain a set of significant coding vectors of the characteristic compensation excited state of the tracking device posture change timing pattern; The set of the tracking device posture change temporal pattern feature compensation excitation state saliency coding vectors and the tracking device posture change temporal pattern set ground state feature aggregation coding vector are input into the feature gain superposition network to obtain the tracking device global dynamic collaborative semantic aggregation coding vector.

8. The neurosurgery navigation system based on augmented reality technology according to claim 7, characterized in that: The operation stage estimation unit includes: inputting the global dynamic collaborative semantic aggregation coding vector of the tracking instrument into a classifier-based operation stage recognition model to obtain the currently estimated operation stage.

9. A neurosurgery navigation method based on augmented reality technology, characterized in that: include: Acquiring marker tracking data collected by an optical tracking system; Based on the landmark tracking data, a real-time position and type list of the tracked device in the current field of view is generated; Inputting the real-time position and type list of the tracked instrument in the current field of view into a surgical stage recognition model to obtain a current estimated surgical stage; extracting a target visualization parameter set from a context-visualization rule database based on the currently estimated surgical stage; Comparing the target visualization parameter set with the visualization parameter set at the current moment and combining the smooth transition parameter to calculate the visualization parameter set used for rendering the next frame; Rendering is performed based on the visualization parameter set used for the next frame rendering to obtain an AR navigation view.

10. The neurosurgery navigation method based on augmented reality technology according to claim 9, characterized in that: Inputting the real-time position and type list of the tracked instrument in the current field of view into the surgical stage recognition model to obtain the current estimated surgical stage, including: Grouping the real-time position and type list of the tracked instrument within the current field of view based on type to obtain a set of real-time position time series of the tracked instrument; Extracting device posture change timing pattern features from each tracking device real-time posture time series in the set of tracking device real-time posture time series to obtain a set of tracking device posture change timing pattern feature encoding vectors; Performing a posture temporal collaborative aggregation analysis on a set of the tracking device posture change temporal pattern feature coding vectors to obtain a global dynamic collaborative semantic aggregation coding vector of the tracking device; The currently estimated surgical stage is determined based on the global dynamic collaborative semantic aggregation coding vector of the tracking instrument.

Citation Information

Patent Citations

  • Abnormality diagnosis method for electric energy metering device

    CN120524332A

  • Steel structure intelligent monitoring and adjusting method and system

    CN120561804A

  • Abnormal traffic detection method based on time sequence

    CN120614168A

Cited By

  • Abnormality diagnosis method for electric energy metering device

    CN120524332A