Urinary surgery training method and system based on augmented reality simulation, electronic equipment and storage medium

By constructing virtual surgical scenarios using augmented reality technology, integrating digital twin models of surgical instruments, and dynamically responding in a real environment, the problem of the separation between the virtual and real worlds in virtual reality training is solved, enabling efficient and safe training in urological surgical skills.

CN121884656AInactive Publication Date: 2026-04-17TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
Filing Date
2026-01-06
Publication Date
2026-04-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing urological surgical training systems suffer from a disconnect between the virtual and real worlds in virtual reality, limited interaction methods, and a lack of high-fidelity feedback and data loops, making it difficult to achieve efficient and safe skills training.

Method used

By constructing a virtual surgical scenario based on augmented reality, integrating real-time interactive digital twin models of surgical instruments, driving them to respond dynamically in the virtual environment, and overlaying them onto the real training environment, combined with multimodal feedback and data acquisition, a closed-loop training experience of operation-perception-feedback is formed.

Benefits of technology

It achieves seamless integration of virtual surgical scenarios with real environments, enhances the immersion of training and the efficiency of skill transfer, provides high-fidelity tactile and visual feedback, supports full-process digital data collection and evaluation, and improves the systematicness and accuracy of training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884656A_ABST
    Figure CN121884656A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical simulation training and surgical education, and discloses a urinary surgery training method and system based on augmented reality simulation, electronic equipment and a storage medium. The method comprises the following steps: S1, constructing a virtual operation scene according to a preset three-dimensional anatomical model of the urinary system; s2, integrating a real-time interactive surgical instrument digital twinborn model in the virtual surgical scene; s3, according to a user operation input signal, driving the surgical instrument digital twin model to dynamically respond in the virtual surgical scene; s4, superposing the dynamically responded virtual operation process into a real training environment visual field through augmented reality display equipment; and S5, generating corresponding operation track data according to continuous operation behaviors of the user in the superposed visual field, and storing the operation track data in a training database. According to the method, the defect of scene splitting in traditional virtual reality training can be effectively overcome through high-precision space registration and real-time rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical simulation training and surgical education technology, specifically to a urological surgical method, system, electronic device, and storage medium based on augmented reality simulation. Background Technology

[0002] Urological surgeries, such as transurethral resection of the prostate (TURP) and percutaneous nephrolithotomy, place extremely high demands on the surgeon's precision, spatial awareness, and hand-eye coordination due to the confined operating space, complex anatomical structures, and proximity to important blood vessels and nerves. Traditional apprenticeship-style clinical training models have inherent drawbacks, including long training cycles, limited case resources, high risks, and difficulty in standardization. With the widespread adoption of minimally invasive surgical techniques, how to efficiently, safely, and reproducibly cultivate surgeons' core surgical skills in non-clinical settings has become an urgent need in medical education.

[0003] To overcome the limitations of traditional training, various surgical simulation training technologies have emerged. Early training primarily relied on box simulators and excised animal organs for basic skills, but these methods had significant shortcomings in anatomical realism, tissue response fidelity, and the diversity of training scenarios. With the development of computer technology, virtual reality (VR)-based surgical simulators have gradually become mainstream. These systems construct three-dimensional virtual anatomical environments, allowing trainees to practice risk-free procedures in a completely virtual space and providing quantitative feedback on aspects such as operation time and path error. However, existing VR simulators typically completely isolate users from the real environment; the operational scenarios, instruments, and tactile sensations are disconnected from the actual operating room environment, limiting the effective transfer of skills to real surgical scenarios. Furthermore, most systems focus on visual simulation, leaving considerable room for improvement in the realism of force and tactile feedback and the physical fidelity of instrument-tissue interaction.

[0004] In recent years, augmented reality (AR) technology has offered new insights for surgical training due to its ability to overlay virtual information onto the real world. Existing research has attempted to project simple 3D models onto physical models to aid anatomical understanding. However, current AR-based training solutions are mostly in the demonstration or early application stages, and generally suffer from the following problems: insufficient geometric registration accuracy between virtual content and the real environment, leading to "virtual-real misalignment"; limited interaction methods, lacking high-fidelity digital twins and real-time dynamic responses to the operation of real surgical instruments; and systems that primarily display information in a one-way manner, failing to form a complete data loop covering operation acquisition, process reproduction, and intelligent evaluation, thus limiting their effectiveness as a systematic and intelligent training tool.

[0005] Therefore, there is an urgent need for a new urological surgical training method that can deeply integrate high-fidelity virtual simulation, natural and intuitive augmented reality presentation, precise physical interactive feedback, and full-process data-driven intelligent assessment to fill the gap in existing technology and achieve more efficient, safer, and quantifiable surgical skills training. Summary of the Invention

[0006] To address the above technical problems, this invention provides a urological surgery training method based on augmented reality simulation, comprising the following steps:

[0007] S1. Construct a virtual surgical scene based on a pre-set three-dimensional anatomical model of the urinary system;

[0008] S2. In the virtual surgical scenario, integrate a real-time interactive digital twin model of surgical instruments;

[0009] S3. Drive the digital twin model of the surgical instrument to dynamically respond in the virtual surgical scene according to the user's operation input signal;

[0010] S4. The virtual surgical procedure after dynamic response is superimposed onto the field of view of the real training environment through an augmented reality display device;

[0011] S5. Based on the user's continuous operation behavior in the overlay field of view, generate corresponding operation trajectory data and store it in the training database.

[0012] Preferably, S2 includes:

[0013] Obtain key parameters of physical surgical instruments, including: geometric parameters, degrees of freedom of motion, and force feedback characteristics;

[0014] A dynamic simulation model of the surgical instrument was established based on the aforementioned key parameters;

[0015] The dynamic simulation model is registered with the spatial coordinate system in the virtual surgical scene to integrate a real-time interactive digital twin model of surgical instruments.

[0016] Preferably, the methods for acquiring user operation input signals include:

[0017] The movement and posture data of the trainee's upper limbs and hands are collected using wearable motion capture devices;

[0018] Based on the motion posture data, the estimated values ​​of operation direction, speed, and contact force are analyzed;

[0019] The parsing results are converted into operational input signals that can be recognized in the virtual environment.

[0020] Preferably, the step of driving the digital twin model of the surgical instrument to perform dynamic response includes:

[0021] The positional changes of the instrument in virtual space are calculated based on the received operational input signals;

[0022] Determine whether the pose change causes a spatial collision with the virtual anatomical structure;

[0023] If a collision occurs, the local organ morphology is updated according to the preset tissue deformation model, and the resistance signal is fed back to the user.

[0024] Preferably, S4 includes:

[0025] The user's viewpoint position and orientation are determined by the spatial positioning module built into the head-mounted display device;

[0026] Real-time rendering of perspective images of the virtual surgical procedure based on viewpoint information;

[0027] The perspective image is overlaid as a transparent layer onto the actual operating table view captured by the camera.

[0028] Preferably, the superimposed field of view further includes auxiliary guiding elements, wherein the method for generating the auxiliary guiding elements includes:

[0029] Identify key anatomical landmarks based on the current stage of surgery;

[0030] Retrieve the standard operation path template for this stage from the training database;

[0031] Based on the deviation between the standard operation path and the user's actual operation trajectory, visual prompts or highlighted areas are generated and superimposed on the field of view.

[0032] Preferably, S5 includes:

[0033] Real-time recording of the spatial coordinate sequence of the instrument tip during user operation;

[0034] Simultaneously collect operation time points, estimated applied force values, and angle changes in viewing perspective to generate multi-dimensional time-series data;

[0035] Multidimensional time-series data is encapsulated into structured operation trajectory data packets using a unified timestamp.

[0036] The present invention also provides a urological surgery training system based on augmented reality simulation. The system is used to implement the above method and includes: a construction module, an integration module, a driving module, an overlay module, and a generation module.

[0037] The construction module is used to construct a virtual surgical scene based on a preset three-dimensional anatomical model of the urinary system;

[0038] The integration module is used to integrate real-time interactive digital twin models of surgical instruments in the virtual surgical scenario;

[0039] The driving module is used to drive the digital twin model of the surgical instrument to dynamically respond in the virtual surgical scene according to the user's operation input signal;

[0040] The overlay module is used to overlay the virtual surgical procedure after dynamic response onto the field of view of the real training environment through an augmented reality display device;

[0041] The generation module is used to generate corresponding operation trajectory data based on the user's continuous operation behavior in the superimposed field of view and store it in the training database.

[0042] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method.

[0043] The present invention also provides a computer-readable storage medium storing a computer program that, when executed, implements the above-described method.

[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0045] This invention enables seamless and stable overlay of virtual surgical scenarios with realistic physical interaction characteristics onto a real training environment through high-precision spatial registration and real-time rendering. This effectively overcomes the drawbacks of fragmented scenarios in traditional virtual reality training, significantly improving the immersion of training and the efficiency of clinical skill transfer. Simultaneously, by integrating high-fidelity instrument digital twins, physics engine simulation, and multimodal feedback, the system achieves realistic simulation of tactile feedback, tissue deformation visuals, and force-tactile feedback, constructing a closed-loop training experience of "operation-perception-feedback," significantly enhancing the training effect of hand-eye coordination and fine force control skills. Furthermore, the solution achieves full digitization of the training process, automatically collecting and structurally storing multidimensional operational data. This supports data-driven real-time guidance, objective quantitative assessment of operational skills, and long-term personalized training planning, greatly improving the systematicness, accuracy, and scientific nature of surgical skills training. Attached Figure Description

[0046] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1This is a schematic diagram of the method flow according to an embodiment of the present invention;

[0048] Figure 2 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention.

[0049] Explanation of reference numerals in the attached figures:

[0050] 1010, Processor; 1020, Memory; 1030, Input / Output Interface; 1040, Communication Interface; 1050, Bus. Detailed Implementation

[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0053] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0054] Example 1

[0055] As can be seen from the background technology, with the popularization of minimally invasive surgical techniques, how to efficiently, safely and repeatedly train surgeons' core surgical skills in non-clinical environments has become an urgent need in the field of medical education.

[0056] Based on this, embodiments of the present invention provide a method for training urological surgery based on augmented reality simulation, comprising the following steps:

[0057] S1. Construct a virtual surgical scene based on a pre-set three-dimensional anatomical model of the urinary system.

[0058] The model in this step can be established by three-dimensional reconstruction, segmentation and annotation of a large amount of real medical image data (such as CT and MRI sequences). Alternatively, it can be combined with publicly available anatomical atlas databases for parametric modeling and personalized correction to ensure that the model accurately reflects the anatomical relationships and variations of urinary organs such as the kidneys, ureters, bladder, and prostate, as well as important blood vessels and nerves around them.

[0059] The process of constructing a virtual surgical scene uses the 3D anatomical model as the core spatial carrier, further integrating surgical environment elements, a physics engine, and a rendering pipeline to form a simulated world in which trainees can immerse themselves for operational training. Specifically, the 3D anatomical model of the urinary system is first placed in a virtual 3D spatial coordinate system, typically defined as the world coordinate system, serving as a reference benchmark for the spatial position and orientation of all virtual objects. Based on this, necessary virtual elements are added to the scene according to the actual operational environment of typical urological surgeries (such as transurethral resection of the prostate and percutaneous nephrolithotomy). These elements include, but are not limited to: the shadowless lighting effects simulating the operating table and its drapes, particle systems for simulating the flow of irrigation fluid, and reference frames representing patient positioning. The scene construction must fully consider visual realism; therefore, physically accurate materials and shaders must be applied to the model and environmental elements. For example, materials with appropriate reflective and diffuse properties are given to organ surfaces, different colors and transparency are assigned to different tissues such as muscles and fat, and global illumination and local shadows are configured to enhance the sense of depth and three-dimensionality.

[0060] During the scene construction phase, it is necessary to pre-integrate or define the logic and data interfaces related to interaction. For example, interactive areas (such as surgical incision locations and lesion areas) and non-collision-prone areas (such as important blood vessels and nerve bundles) need to be marked on the 3D anatomical model. Corresponding physical property parameters (such as elastic modulus, coefficient of friction, and fracture threshold) should be assigned to different tissue types (such as mucosa, muscle layer, and bone). These properties will serve as the basis for collision detection, tissue deformation simulation, and force feedback calculation in subsequent steps. At the same time, the scene construction must ensure compatibility with the rendering pipeline of the augmented reality display device. That is, the rendering perspective, perspective relationship, and resolution of the virtual scene can be adjusted in real time according to the user's viewpoint and ultimately superimposed onto the real field of view with the correct spatial registration relationship.

[0061] S2. In the virtual surgical scenario, integrate a real-time interactive digital twin model of surgical instruments.

[0062] The integration of digital twin models is a core component in constructing high-fidelity, interactive virtual training scenarios. Its goal is to accurately reproduce the geometry, movement, and tactile sensations of real surgical instruments within a virtual environment, thereby providing trainees with a near-realistic hand-operation experience.

[0063] The integration process begins with a deep digital characterization of the target physical surgical instruments. This requires acquiring key instrument parameters, encompassing multiple dimensions including geometry, kinematics, and mechanics. Geometric parameters are obtained through high-precision 3D scanning or parametric modeling based on engineering drawings, accurately describing the instrument's overall shape, local features (such as jaw teeth and electrode shapes), and key dimensions. The degrees of freedom parameters define the range of motion, interrelationships, and constraints of each moving part of the instrument (such as the bending of the clamp bar, the opening and closing of the jaws, and the rotation of the endoscope). Force feedback characteristics quantify key sensory parameters, involving the resistance characteristics, elastic feedback, and vibration sensations exhibited by the instrument when in contact with different tissues. These data can be obtained through calibration experiments on simulated tissues or standard materials using specialized force measuring devices and converted into mechanical model parameters related to displacement and velocity.

[0064] Based on the acquired key parameters, the next step is to establish a dynamic simulation model of the surgical instrument. This model is a computational entity that simulates the motion and force state of the instrument in a virtual environment according to physical laws (mainly Newtonian mechanics and rigid / flexible body dynamics). During modeling, the instrument is abstracted as a multibody system composed of multiple rigid or flexible bodies connected by joints and constraints. The mass distribution and collision volume of each component are defined using the acquired geometric parameters, and the joint types (such as rotation and translation) and their motion limits are set according to the degrees of freedom parameters. Force feedback characteristics are integrated into the contact mechanics model. For example, a nonlinear spring-damping system based on penetration depth is used to simulate the indentation and rebound when the instrument contacts soft tissue, or a Coulomb friction model is used to simulate the sliding resistance of the instrument within the lumen. This process is typically achieved using a mature physics engine (such as PhysX or Bullet). The engine is responsible for calculating the new position, new posture, and generated reaction forces of each component of the instrument in each simulation step, based on the user-input operation commands (converted into forces or displacements applied to the model).

[0065] To ensure the correct operation of a digital twin model in a virtual surgical scenario, it must be uniformly registered with the virtual scene's spatial coordinate system. The virtual surgical scene possesses its own world coordinate system, upon which all anatomical structures and environmental elements are positioned. During integration, a clear "anchor point" or "handle" coordinate system needs to be defined for the digital twin model. This coordinate system typically corresponds to the actual gripping point of the instrument held by the trainee or the reference point of the motion capture device. A calibration procedure establishes the transformation relationship (including rotation and translation) between this handle coordinate system and the virtual world coordinate system. When the user operates a real input device (such as a force feedback robotic arm or wearable sensors), their motion data is captured in real time and converted to the virtual world coordinate system, thereby driving the digital twin model to perform corresponding pose changes. Simultaneously, the collision detection and force interactions between instruments and virtual anatomical structures calculated by the physics engine are also performed entirely within this unified spatial framework, ensuring that the visual movement of the instrument model, the collision deformation with organs, and the force signals fed back to the user are completely consistent in spatial logic. Ultimately, this digital twin model of surgical instruments, with precisely defined parameters, dynamic simulation capabilities, and alignment with the virtual scene space, was seamlessly integrated into the pre-built virtual surgical scene. It became a core interactive entity that could be driven in real time and had realistic physical interactions with the virtual anatomical environment, laying a solid foundation for responding to user operations and simulating the surgical process in subsequent steps.

[0066] S3. Based on the user's input signal, drive the digital twin model of the surgical instrument to dynamically respond in the virtual surgical scene.

[0067] After integrating the high-fidelity digital twin model of surgical instruments into the virtual scene, the core task of this step is to respond to the user's operational intentions in real time, driving the model to perform realistic and physically consistent movements and interactions in the complex virtual anatomical environment. This dynamic response process constitutes the core driving link of the simulation training interaction closed loop. It is not a simple visual model displacement, but a complex calculation and rendering process that integrates high-precision pose calculation, real-time physical collision detection, biomechanical tissue deformation simulation, and multimodal feedback generation.

[0068] The dynamic response is triggered by the acquisition and parsing of user input signals. Through motion capture devices worn on the trainee's upper limbs and hands, raw motion posture data, such as joint angles, limb positions, and orientations, are continuously collected. This data undergoes filtering and fusion processing to eliminate noise and establish a stable human kinematic chain. Further analytical algorithms extract operational instructions with clear surgical semantics from this data: including the expected direction of instrument movement, instantaneous velocity, and estimated contact force estimated through indirect measurements of muscle activity electrical signals or specific pressure sensors. All this information is encapsulated and converted into a standardized set of operational input signals that the virtual environment engine can recognize. This signal not only contains spatial displacement commands but also carries the user's intention to apply "force," serving as a key input to drive the digital twin model for dynamic simulation, rather than merely kinematic animation.

[0069] Upon receiving the operation input signal, the digital twin model is first driven to perform pose calculations. This model is typically expressed as a multibody dynamics equation to solve for the motion state of the device in virtual space. Its general form can be expressed as:

[0070]

[0071] in, Represents the system's inertia matrix; and These are the first and second derivatives of the generalized coordinate q with respect to time, respectively; Represents the Coriolis force and centripetal force terms; Represents the gravity term; This represents the vector of active forces / torques acting on the system's generalized coordinates. Indicates contact force The contribution to the generalized coordinate system is the virtual work done by the contact force on the generalized coordinate system.

[0072] This process relies on the instrument dynamics simulation model and physics engine established during the integration phase. The direction and velocity information contained in the input signals are converted into forces or torques applied to specific "handles" or drive points on the model. The physics engine then uses real-time numerical integration (such as the Euler method or the Runge-Kutta method) to calculate the model's new position, new posture (collectively referred to as "pose"), and the state of each moving part in the next simulation time step, based on this driving force, the model's own mass, moment of inertia, and global physical parameters such as gravity and damping defined in the virtual environment. This calculation ensures that the motion of the instrument model follows Newton's laws of motion. For example, a rapid push will cause the instrument to slightly overshoot due to inertia, and when the jaws are released, it will automatically rebound due to its internal spring model, thus presenting a realistic sense of motion.

[0073] Simultaneously and subsequently, the engine compares the expected new pose of the instrument's digital twin model (typically represented by its simplified collision mesh) with a precise 3D model of all anatomical structures in the virtual surgical scene. The core of collision detection is determining geometric penetration. For a point P (such as the instrument tip) and a triangular facet (the surface of an anatomical structure), a preliminary judgment can be made by calculating the signed distance d from point P to the plane containing the triangular facet. If point P is within the projection area of ​​the triangular facet and d is less than zero (or a small positive threshold ε), it is considered a potential collision. A more precise penetration depth d_collision can be calculated using the formula:

[0074]

[0075] Where P is the spatial coordinate vector of the point to be tested on the instrument; S0 is the coordinate vector of any vertex on the triangular facet; n is the unit normal vector of the triangular facet, pointing outwards from the object; · represents the dot product operation of vectors; if d_collision<0, it means that point P penetrates into the interior of the triangular facet, with a penetration depth of |d_collision|.

[0076] If the detection result is "no collision occurred", the calculated new pose is directly assigned to the instrument model to complete a smooth movement in the air or in the cavity, and may generate corresponding visual eddies or particle effects based on the movement speed and the medium model (such as simulated irrigation fluid).

[0077] However, if the collision detection determines that the instrument has made spatial contact with the virtual anatomical structure, a multiphysics response process is immediately initiated. First, based on a pre-set tissue deformation model with different tissue biomechanical properties, the morphology of the local organ at the point of contact is updated in real time. For example, a mass-spring system or finite element model is used to simulate the elastic deformation, indentation, and even stretching of soft tissue under instrument compression; for more complex cutting or electrocautery operations, the model mesh may be dynamically modified by combining the tissue fracture threshold or heat conduction model to simulate the morphological changes of tissue cutting and carbonization. This deformation process is the core of the visual response, making the virtual organ no longer a rigid body, but a living tissue capable of interacting with the instrument and producing realistic deformation.

[0078] Simultaneously, in sync with the deformation visual feedback, a corresponding tactile resistance signal is generated and fed back to the user. The physics engine calculates the reaction force and torque acting on the device based on the collision depth, the normal direction of the contact surface, the preset physical properties of the tissue (such as elastic modulus and viscosity coefficient), and the user's movement speed. This force signal is not simply constant but dynamically changing: touching mucous membranes may result in slight elastic resistance, the resistance increases linearly when force is applied, and touching bone presents a hard, bottom-feeling sensation. This calculated force feedback signal is converted into a real mechanical force acting on the trainee's hand through an integrated force feedback device (such as a high-fidelity force feedback robotic arm or professional tactile gloves), simulating the tactile sensation during device operation, completing the closed loop from virtual interaction to real tactile sensation.

[0079] The entire dynamic response process cycles at an extremely high frequency (typically requiring no less than 90Hz to maintain immersion), ensuring a high degree of synchronization and consistency between the user's input, the visual movement of the instrument model, tissue deformation, and the force feedback from the hand. Through this series of precisely linked steps, the system successfully transforms the trainee's physical actions into an interactive process with realistic physical meaning and visual effects in a virtual surgical scenario. This lays the foundation for the dynamic and interactive core content of subsequently overlaying this realistic process in the augmented reality field and recording and analyzing the operational data.

[0080] S4. The virtual surgical procedure after dynamic response is superimposed onto the field of view of the real training environment through an augmented reality display device.

[0081] First, it is necessary to accurately perceive the user's viewing angle in the real physical space. This task is accomplished by the high-precision spatial positioning module built into the augmented reality display device, typically a head-mounted display (HMD). This module may integrate an inertial measurement unit (IMU), an indoor optical or infrared tracking system, and a simultaneous localization and mapping (SLAM) algorithm. Through the collaborative work of these sensors and algorithms, the system can calculate in real time the three-dimensional position vector P_view and orientation quaternion Q_view (or equivalent rotation matrix R_view) of the user's head (i.e., the viewpoint) in the real-world coordinate system. This P_view and Q_view are the bridge connecting the virtual and the real world; they define the position and orientation of the "camera" in the virtual world to ensure that the virtual image rendered by the "camera" is completely consistent with the real world seen by the user through the display device in terms of perspective geometry.

[0082] Based on real-time acquired viewpoint information, the system initiates a high-performance real-time rendering pipeline to generate perspective images of the virtual surgical procedure. This rendering process takes as input the entire virtual surgical scene that has undergone dynamic responses in step S3—including the digital twin model of surgical instruments whose pose has been updated according to dynamic equations, the anatomical structures of the urinary system whose morphology has changed according to tissue deformation models, and virtual surgical environment elements (such as shadowless lighting effects and liquid particles). The rendering engine sets up the virtual camera according to P_view and Q_view and calculates the corresponding view matrix V and perspective projection matrix P. The view matrix V, determined by P_view and R_view, is used to transform all vertices in the virtual scene from the world coordinate system to the camera's viewing coordinate system. The perspective projection matrix P simulates the imaging characteristics of the human eye or optical lens, projecting three-dimensional coordinates onto a two-dimensional imaging plane and producing a perspective effect where objects appear larger when closer and smaller when farther away. The rendering engine then performs lighting calculations, shadow generation, material shading, and other steps, ultimately outputting a frame of virtual scene color image I_virtual that perfectly matches the current user's viewpoint, possessing correct stereoscopic sense and depth, and a corresponding depth buffer image D_virtual. This process requires an extremely high frame rate (typically ≥90Hz) to match human visual perception and head movements, preventing dizziness caused by screen lag.

[0083] The next crucial step in achieving virtual-real fusion is overlaying the rendered virtual image (I_virtual) onto the real-world field of view. Head-mounted displays typically feature outward-facing cameras that continuously capture images of the real operating table and training environment in front of the user, resulting in a real-world image (I_real). Simple layer overlay can cause virtual objects to completely obscure real objects behind them, disrupting immersion. Therefore, the system employs depth-based intelligent fusion technology. Using the D_virtual generated during rendering, the system can accurately determine the depth information of each pixel in the virtual scene. Simultaneously, through the device's depth sensor or the depth map D_real estimated by computer vision algorithms from I_real, a rough depth information of the real scene can be obtained. During pixel-level fusion, the algorithm compares the D_virtual and D_real values ​​at the same pixel location. If the depth value of a virtual object pixel is less than (i.e., closer to the observer) the depth value of the corresponding location in the real scene, the virtual pixel is retained; otherwise, the real scene pixel should be placed in front of the virtual object, thus retaining the real pixel or making the virtual pixel semi-transparent. Furthermore, I_virtual is often processed as a transparent layer with an alpha channel. By applying anti-aliasing and appropriate feathering to the edges of virtual objects, the transition between them and the real background is made more natural, eliminating harsh boundaries. Finally, the composite image I_composite is sent to the near-eye display of the head-mounted display device, allowing the user to perceive virtual surgical instruments and organs as if they were truly present on the operating table in front of them.

[0084] Furthermore, to enhance training effectiveness, auxiliary guidance elements are dynamically integrated into the overlaid composite field of view. Based on the surgical stage set in the current training (e.g., "locating the ureteral orifice"), the system automatically identifies the coordinate set L_key of key anatomical landmarks in the virtual 3D anatomical model. Simultaneously, it retrieves an expert-level standard operating path template for that stage from a pre-built training database. This template may be represented as a sequence of spatial coordinate points ordered by time: Path_standard. The system tracks the coordinate sequence Path_user of the instrument tip actually manipulated by the user in the overlaid field of view in real time. By calculating the deviation vector δ between Path_user and Path_standard in position, direction, or velocity, the system generates intuitive visual prompts. For example, it dynamically draws an arrow icon pointing from the current instrument position to the next standard path point in the rendering layer, its direction and length determined by the deviation δ; or it highlights the tissue surface where the user deviates from the standard area. These auxiliary guidance elements, I_guidance, are rendered and blended into I_real along with I_virtual as another graphical layer, providing trainees with real-time navigation and error correction feedback, thereby elevating the display value of augmented reality from simple scene overlay to the level of intelligent guidance.

[0085] S5. Based on the user's continuous operation behavior in the overlay field of view, generate corresponding operation trajectory data and store it in the training database.

[0086] First, the spatial coordinate sequence P_tip(t) of the tip of the instrument's digital twin model in the virtual world coordinate system is recorded in real time and with high precision during the user's operation. This sequence captures the movement path of the instrument at a high sampling rate (typically ≥100Hz). Simultaneously, the system adds a timestamp t accurate to the millisecond level and associates it with the estimated force F_estimate(t) exerted by the user on the instrument, obtained in step S3, and the user's viewpoint change angle θ_view(t) tracked through a head-mounted device. These data together constitute a multi-dimensional temporal data stream with time as the axis, fully characterizing the dynamic features of an operation in space, mechanics, and viewpoint.

[0087] Subsequently, the system processes and encapsulates this heterogeneous, high-speed generated raw data in real time. The collected data streams, including spatial coordinates, force estimates, and viewpoints, are strictly aligned and synchronized using a unified timestamp `t` to ensure accurate correspondence between various data types at the same moment. Next, the system encapsulates this aligned data according to a preset structured format, generating an independent "operation trajectory data package." This data package not only contains the aforementioned core time-series data but also typically includes metadata such as trainer identifiers, training task types, the version of the virtual anatomical model used, and the absolute times of training start and end. This structured encapsulation method allows each independent surgical procedure to be stored completely and systematically as a digital archive that can be efficiently parsed and retrieved by a computer.

[0088] Ultimately, these generated structured operation trajectory data packets are transmitted and stored in a central or distributed training database. The database not only serves as a repository for massive amounts of operation records but also enables subsequent in-depth analysis. For example, by accessing historical data in the database, the system can compare the current user's operation trajectory with expert standard templates or the user's previous training data to quantify their operational accuracy, speed, stability, and mechanical control level. Furthermore, these accumulated data resources can also be used for automatic skill level classification based on machine learning, typical error pattern recognition, and even to provide data-driven decision support for generating personalized advanced training programs. This truly realizes the digitalization, measurability, and intelligence of the training process, completing a full training loop from real-time interactive simulation to data-driven evaluation.

[0089] This invention's technical solution seamlessly and stably overlays virtual surgical scenes with realistic physical interaction characteristics onto a real training environment through high-precision spatial registration and real-time rendering. This effectively overcomes the drawbacks of fragmented scenes in traditional virtual reality training, significantly improving the immersion of training and the efficiency of clinical skill transfer. Simultaneously, by integrating high-fidelity instrument digital twins, physics engine simulation, and multimodal feedback, the system achieves realistic simulation of tactile feedback, tissue deformation visuals, and force-tactile feedback, constructing a closed-loop training experience of "operation-perception-feedback," significantly enhancing the training effect of hand-eye coordination and fine force control skills. Furthermore, the solution achieves full digitization of the training process, automatically collecting and structurally storing multidimensional operational data. This supports data-driven real-time guidance, objective quantitative assessment of operational skills, and long-term personalized training planning, greatly improving the systematicness, accuracy, and scientific nature of surgical skills training.

[0090] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0091] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, it should be understood that the sequence number of each step in the above embodiments does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. The actions or steps recorded in the claims can be performed in a different order than that in the above embodiments and can still achieve the desired result. In addition, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0092] Example 2

[0093] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides a urological surgery training system based on augmented reality simulation, including: a construction module, an integration module, a driving module, an overlay module, and a generation module.

[0094] The following will describe in detail, with reference to this embodiment, how the present invention solves the technical problems in practical work.

[0095] First, a virtual surgical scene is constructed using the building module based on a pre-set three-dimensional anatomical model of the urinary system.

[0096] The model in this embodiment can be established by three-dimensional reconstruction, segmentation and annotation of a large amount of real medical image data (such as CT and MRI sequences). It can also be combined with publicly available anatomical atlas databases for parametric modeling and personalized correction to ensure that the model accurately reflects the anatomical relationships and variations of urinary organs such as the kidneys, ureters, bladders, and prostates and their surrounding important blood vessels and nerves.

[0097] The process of constructing a virtual surgical scene uses the 3D anatomical model as the core spatial carrier, further integrating surgical environment elements, a physics engine, and a rendering pipeline to form a simulated world in which trainees can immerse themselves for operational training. Specifically, the 3D anatomical model of the urinary system is first placed in a virtual 3D spatial coordinate system, typically defined as the world coordinate system, serving as a reference benchmark for the spatial position and orientation of all virtual objects. Based on this, necessary virtual elements are added to the scene according to the actual operational environment of typical urological surgeries (such as transurethral resection of the prostate and percutaneous nephrolithotomy). These elements include, but are not limited to: the shadowless lighting effects simulating the operating table and its drapes, particle systems for simulating the flow of irrigation fluid, and reference frames representing patient positioning. The scene construction must fully consider visual realism; therefore, physically accurate materials and shaders must be applied to the model and environmental elements. For example, materials with appropriate reflective and diffuse properties are given to organ surfaces, different colors and transparency are assigned to different tissues such as muscles and fat, and global illumination and local shadows are configured to enhance the sense of depth and three-dimensionality.

[0098] During the scene construction phase, it is necessary to pre-integrate or define the logic and data interfaces related to interaction. For example, interactive areas (such as surgical incision locations and lesion areas) and non-collision-prone areas (such as important blood vessels and nerve bundles) need to be marked on the 3D anatomical model. Corresponding physical property parameters (such as elastic modulus, coefficient of friction, and fracture threshold) should be assigned to different tissue types (such as mucosa, muscle layer, and bone). These properties will serve as the basis for collision detection, tissue deformation simulation, and force feedback calculation in subsequent steps. At the same time, the scene construction must ensure compatibility with the rendering pipeline of the augmented reality display device. That is, the rendering perspective, perspective relationship, and resolution of the virtual scene can be adjusted in real time according to the user's viewpoint and ultimately superimposed onto the real field of view with the correct spatial registration relationship.

[0099] The integration module then integrates real-time interactive digital twin models of surgical instruments into the virtual surgical scenario.

[0100] The integration of digital twin models is a core component in constructing high-fidelity, interactive virtual training scenarios. Its goal is to accurately reproduce the geometry, movement, and tactile sensations of real surgical instruments within a virtual environment, thereby providing trainees with a near-realistic hand-operation experience.

[0101] The integration process begins with a deep digital characterization of the target physical surgical instruments. This requires acquiring key instrument parameters, encompassing multiple dimensions including geometry, kinematics, and mechanics. Geometric parameters are obtained through high-precision 3D scanning or parametric modeling based on engineering drawings, accurately describing the instrument's overall shape, local features (such as jaw teeth and electrode shapes), and key dimensions. The degrees of freedom parameters define the range of motion, interrelationships, and constraints of each moving part of the instrument (such as the bending of the clamp bar, the opening and closing of the jaws, and the rotation of the endoscope). Force feedback characteristics quantify key sensory parameters, involving the resistance characteristics, elastic feedback, and vibration sensations exhibited by the instrument when in contact with different tissues. These data can be obtained through calibration experiments on simulated tissues or standard materials using specialized force measuring devices and converted into mechanical model parameters related to displacement and velocity.

[0102] Based on the acquired key parameters, the next step is to establish a dynamic simulation model of the surgical instrument. This model is a computational entity that simulates the motion and force state of the instrument in a virtual environment according to physical laws (mainly Newtonian mechanics and rigid / flexible body dynamics). During modeling, the instrument is abstracted as a multibody system composed of multiple rigid or flexible bodies connected by joints and constraints. The mass distribution and collision volume of each component are defined using the acquired geometric parameters, and the joint types (such as rotation and translation) and their motion limits are set according to the degrees of freedom parameters. Force feedback characteristics are integrated into the contact mechanics model. For example, a nonlinear spring-damping system based on penetration depth is used to simulate the indentation and rebound when the instrument contacts soft tissue, or a Coulomb friction model is used to simulate the sliding resistance of the instrument within the lumen. This process is typically achieved using a mature physics engine (such as PhysX or Bullet). The engine is responsible for calculating the new position, new posture, and generated reaction forces of each component of the instrument in each simulation step, based on the user-input operation commands (converted into forces or displacements applied to the model).

[0103] To ensure the correct operation of a digital twin model in a virtual surgical scenario, it must be uniformly registered with the virtual scene's spatial coordinate system. The virtual surgical scene possesses its own world coordinate system, upon which all anatomical structures and environmental elements are positioned. During integration, a clear "anchor point" or "handle" coordinate system needs to be defined for the digital twin model. This coordinate system typically corresponds to the actual gripping point of the instrument held by the trainee or the reference point of the motion capture device. A calibration procedure establishes the transformation relationship (including rotation and translation) between this handle coordinate system and the virtual world coordinate system. When the user operates a real input device (such as a force feedback robotic arm or wearable sensors), their motion data is captured in real time and converted to the virtual world coordinate system, thereby driving the digital twin model to perform corresponding pose changes. Simultaneously, the collision detection and force interactions between instruments and virtual anatomical structures calculated by the physics engine are also performed entirely within this unified spatial framework, ensuring that the visual movement of the instrument model, the collision deformation with organs, and the force signals fed back to the user are completely consistent in spatial logic. Ultimately, this digital twin model of surgical instruments, with precisely defined parameters, dynamic simulation capabilities, and alignment with the virtual scene space, was seamlessly integrated into the pre-built virtual surgical scene. It became a core interactive entity that could be driven in real time and had realistic physical interactions with the virtual anatomical environment, laying a solid foundation for responding to user operations and simulating the surgical process in subsequent steps.

[0104] The driving module drives the digital twin model of the surgical instruments to dynamically respond in the virtual surgical scene based on the user's operation input signals.

[0105] After integrating the high-fidelity digital twin model of surgical instruments into the virtual scene, the core task of this step is to respond to the user's operational intentions in real time, driving the model to perform realistic and physically consistent movements and interactions in the complex virtual anatomical environment. This dynamic response process constitutes the core driving link of the simulation training interaction closed loop. It is not a simple visual model displacement, but a complex calculation and rendering process that integrates high-precision pose calculation, real-time physical collision detection, biomechanical tissue deformation simulation, and multimodal feedback generation.

[0106] The dynamic response is triggered by the acquisition and parsing of user input signals. Through motion capture devices worn on the trainee's upper limbs and hands, raw motion posture data, such as joint angles, limb positions, and orientations, are continuously collected. This data undergoes filtering and fusion processing to eliminate noise and establish a stable human kinematic chain. Further analytical algorithms extract operational instructions with clear surgical semantics from this data: including the expected direction of instrument movement, instantaneous velocity, and estimated contact force estimated through indirect measurements of muscle activity electrical signals or specific pressure sensors. All this information is encapsulated and converted into a standardized set of operational input signals that the virtual environment engine can recognize. This signal not only contains spatial displacement commands but also carries the user's intention to apply "force," serving as a key input to drive the digital twin model for dynamic simulation, rather than merely kinematic animation.

[0107] Upon receiving the operation input signal, the digital twin model is first driven to perform pose calculations. This model is typically expressed as a multibody dynamics equation to solve for the motion state of the device in virtual space. Its general form can be expressed as:

[0108]

[0109] in, Represents the system's inertia matrix; and These are the first and second derivatives of the generalized coordinate q with respect to time, respectively; Represents the Coriolis force and centripetal force terms; Represents the gravity term; This represents the vector of active forces / torques acting on the system's generalized coordinates. Indicates contact force The contribution to the generalized coordinate system is the virtual work done by the contact force on the generalized coordinate system.

[0110] This process relies on the instrument dynamics simulation model and physics engine established during the integration phase. The direction and velocity information contained in the input signals are converted into forces or torques applied to specific "handles" or drive points on the model. The physics engine then uses real-time numerical integration (such as the Euler method or the Runge-Kutta method) to calculate the model's new position, new posture (collectively referred to as "pose"), and the state of each moving part in the next simulation time step, based on this driving force, the model's own mass, moment of inertia, and global physical parameters such as gravity and damping defined in the virtual environment. This calculation ensures that the motion of the instrument model follows Newton's laws of motion. For example, a rapid push will cause the instrument to slightly overshoot due to inertia, and when the jaws are released, it will automatically rebound due to its internal spring model, thus presenting a realistic sense of motion.

[0111] Simultaneously and subsequently, the engine compares the expected new pose of the instrument's digital twin model (typically represented by its simplified collision mesh) with a precise 3D model of all anatomical structures in the virtual surgical scene. The core of collision detection is determining geometric penetration. For a point P (such as the instrument tip) and a triangular facet (the surface of an anatomical structure), a preliminary judgment can be made by calculating the signed distance d from point P to the plane containing the triangular facet. If point P is within the projection area of ​​the triangular facet and d is less than zero (or a small positive threshold ε), it is considered a potential collision. A more precise penetration depth d_collision can be calculated using the formula:

[0112]

[0113] Where P is the spatial coordinate vector of the point to be tested on the instrument; S0 is the coordinate vector of any vertex on the triangular facet; n is the unit normal vector of the triangular facet, pointing outwards from the object; · represents the dot product operation of vectors; if d_collision<0, it means that point P penetrates into the interior of the triangular facet, with a penetration depth of |d_collision|.

[0114] If the detection result is "no collision occurred", the calculated new pose is directly assigned to the instrument model to complete a smooth movement in the air or in the cavity, and may generate corresponding visual eddies or particle effects based on the movement speed and the medium model (such as simulated irrigation fluid).

[0115] However, if the collision detection determines that the instrument has made spatial contact with the virtual anatomical structure, a multiphysics response process is immediately initiated. First, based on a pre-set tissue deformation model with different tissue biomechanical properties, the morphology of the local organ at the point of contact is updated in real time. For example, a mass-spring system or finite element model is used to simulate the elastic deformation, indentation, and even stretching of soft tissue under instrument compression; for more complex cutting or electrocautery operations, the model mesh may be dynamically modified by combining the tissue fracture threshold or heat conduction model to simulate the morphological changes of tissue cutting and carbonization. This deformation process is the core of the visual response, making the virtual organ no longer a rigid body, but a living tissue capable of interacting with the instrument and producing realistic deformation.

[0116] Simultaneously, in sync with the deformation visual feedback, a corresponding tactile resistance signal is generated and fed back to the user. The physics engine calculates the reaction force and torque acting on the device based on the collision depth, the normal direction of the contact surface, the preset physical properties of the tissue (such as elastic modulus and viscosity coefficient), and the user's movement speed. This force signal is not simply constant but dynamically changing: touching mucous membranes may result in slight elastic resistance, the resistance increases linearly when force is applied, and touching bone presents a hard, bottom-feeling sensation. This calculated force feedback signal is converted into a real mechanical force acting on the trainee's hand through an integrated force feedback device (such as a high-fidelity force feedback robotic arm or professional tactile gloves), simulating the tactile sensation during device operation, completing the closed loop from virtual interaction to real tactile sensation.

[0117] The entire dynamic response process cycles at an extremely high frequency (typically requiring no less than 90Hz to maintain immersion), ensuring a high degree of synchronization and consistency between the user's input, the visual movement of the instrument model, tissue deformation, and the force feedback from the hand. Through this series of precisely linked steps, the system successfully transforms the trainee's physical actions into an interactive process with realistic physical meaning and visual effects in a virtual surgical scenario. This lays the foundation for the dynamic and interactive core content of subsequently overlaying this realistic process in the augmented reality field and recording and analyzing the operational data.

[0118] The overlay module overlays the virtual surgical procedure after dynamic response onto the real training environment's field of vision using an augmented reality display device.

[0119] First, it is necessary to accurately perceive the user's viewing angle in the real physical space. This task is accomplished by the high-precision spatial positioning module built into the augmented reality display device, typically a head-mounted display (HMD). This module may integrate an inertial measurement unit (IMU), an indoor optical or infrared tracking system, and a simultaneous localization and mapping (SLAM) algorithm. Through the collaborative work of these sensors and algorithms, the system can calculate in real time the three-dimensional position vector P_view and orientation quaternion Q_view (or equivalent rotation matrix R_view) of the user's head (i.e., the viewpoint) in the real-world coordinate system. This P_view and Q_view are the bridge connecting the virtual and the real world; they define the position and orientation of the "camera" in the virtual world to ensure that the virtual image rendered by the "camera" is completely consistent with the real world seen by the user through the display device in terms of perspective geometry.

[0120] Based on real-time acquired viewpoint information, the system initiates a high-performance real-time rendering pipeline to generate perspective images of the virtual surgical procedure. This rendering process takes as input the entire virtual surgical scene that has undergone dynamic responses in step S3—including the digital twin model of surgical instruments whose pose has been updated according to dynamic equations, the anatomical structures of the urinary system whose morphology has changed according to tissue deformation models, and virtual surgical environment elements (such as shadowless lighting effects and liquid particles). The rendering engine sets up the virtual camera according to P_view and Q_view and calculates the corresponding view matrix V and perspective projection matrix P. The view matrix V, determined by P_view and R_view, is used to transform all vertices in the virtual scene from the world coordinate system to the camera's viewing coordinate system. The perspective projection matrix P simulates the imaging characteristics of the human eye or optical lens, projecting three-dimensional coordinates onto a two-dimensional imaging plane and producing a perspective effect where objects appear larger when closer and smaller when farther away. The rendering engine then performs lighting calculations, shadow generation, material shading, and other steps, ultimately outputting a frame of virtual scene color image I_virtual that perfectly matches the current user's viewpoint, possessing correct stereoscopic sense and depth, and a corresponding depth buffer image D_virtual. This process requires an extremely high frame rate (typically ≥90Hz) to match human visual perception and head movements, preventing dizziness caused by screen lag.

[0121] The next crucial step in achieving virtual-real fusion is overlaying the rendered virtual image (I_virtual) onto the real-world field of view. Head-mounted displays typically feature outward-facing cameras that continuously capture images of the real operating table and training environment in front of the user, resulting in a real-world image (I_real). Simple layer overlay can cause virtual objects to completely obscure real objects behind them, disrupting immersion. Therefore, the system employs depth-based intelligent fusion technology. Using the D_virtual generated during rendering, the system can accurately determine the depth information of each pixel in the virtual scene. Simultaneously, through the device's depth sensor or the depth map D_real estimated by computer vision algorithms from I_real, a rough depth information of the real scene can be obtained. During pixel-level fusion, the algorithm compares the D_virtual and D_real values ​​at the same pixel location. If the depth value of a virtual object pixel is less than (i.e., closer to the observer) the depth value of the corresponding location in the real scene, the virtual pixel is retained; otherwise, the real scene pixel should be placed in front of the virtual object, thus retaining the real pixel or making the virtual pixel semi-transparent. Furthermore, I_virtual is often processed as a transparent layer with an alpha channel. By applying anti-aliasing and appropriate feathering to the edges of virtual objects, the transition between them and the real background is made more natural, eliminating harsh boundaries. Finally, the composite image I_composite is sent to the near-eye display of the head-mounted display device, allowing the user to perceive virtual surgical instruments and organs as if they were truly present on the operating table in front of them.

[0122] Furthermore, to enhance training effectiveness, auxiliary guidance elements are dynamically integrated into the overlaid composite field of view. Based on the surgical stage set in the current training (e.g., "locating the ureteral orifice"), the system automatically identifies the coordinate set L_key of key anatomical landmarks in the virtual 3D anatomical model. Simultaneously, it retrieves an expert-level standard operating path template for that stage from a pre-built training database. This template may be represented as a sequence of spatial coordinate points ordered by time: Path_standard. The system tracks the coordinate sequence Path_user of the instrument tip actually manipulated by the user in the overlaid field of view in real time. By calculating the deviation vector δ between Path_user and Path_standard in position, direction, or velocity, the system generates intuitive visual prompts. For example, it dynamically draws an arrow icon pointing from the current instrument position to the next standard path point in the rendering layer, its direction and length determined by the deviation δ; or it highlights the tissue surface where the user deviates from the standard area. These auxiliary guidance elements, I_guidance, are rendered and blended into I_real along with I_virtual as another graphical layer, providing trainees with real-time navigation and error correction feedback, thereby elevating the display value of augmented reality from simple scene overlay to the level of intelligent guidance.

[0123] Finally, the generation module generates corresponding operation trajectory data based on the user's continuous operation behavior in the overlay field of view and stores it in the training database.

[0124] First, the spatial coordinate sequence P_tip(t) of the tip of the instrument's digital twin model in the virtual world coordinate system is recorded in real time and with high precision during the user's operation. This sequence captures the movement path of the instrument at a high sampling rate (typically ≥100Hz). Simultaneously, the system adds a timestamp t accurate to the millisecond level and associates it with the estimated force F_estimate(t) exerted by the user on the instrument, obtained in step S3, and the user's viewpoint change angle θ_view(t) tracked through a head-mounted device. These data together constitute a multi-dimensional temporal data stream with time as the axis, fully characterizing the dynamic features of an operation in space, mechanics, and viewpoint.

[0125] Subsequently, the system processes and encapsulates this heterogeneous, high-speed generated raw data in real time. The collected data streams, including spatial coordinates, force estimates, and viewpoints, are strictly aligned and synchronized using a unified timestamp `t` to ensure accurate correspondence between various data types at the same moment. Next, the system encapsulates this aligned data according to a preset structured format, generating an independent "operation trajectory data package." This data package not only contains the aforementioned core time-series data but also typically includes metadata such as trainer identifiers, training task types, the version of the virtual anatomical model used, and the absolute times of training start and end. This structured encapsulation method allows each independent surgical procedure to be stored completely and systematically as a digital archive that can be efficiently parsed and retrieved by a computer.

[0126] Ultimately, these generated structured operation trajectory data packets are transmitted and stored in a central or distributed training database. The database not only serves as a repository for massive amounts of operation records but also enables subsequent in-depth analysis. For example, by accessing historical data in the database, the system can compare the current user's operation trajectory with expert standard templates or the user's previous training data to quantify their operational accuracy, speed, stability, and mechanical control level. Furthermore, these accumulated data resources can also be used for automatic skill level classification based on machine learning, typical error pattern recognition, and even to provide data-driven decision support for generating personalized advanced training programs. This truly realizes the digitalization, measurability, and intelligence of the training process, completing a full training loop from real-time interactive simulation to data-driven evaluation.

[0127] The system described in the above embodiments is used to implement a corresponding augmented reality simulation-based urological surgery training method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0128] It should be noted that the aforementioned augmented reality simulation-based urological surgery training system is embodied in the form of functional units. The term "module" here can be implemented in software and / or hardware, without specific limitations.

[0129] For example, a "module" can be a software program, hardware circuit, or a combination of both that implements the above functions. Hardware circuits may include application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functions.

[0130] Example 3

[0131] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a urological surgery training method based on augmented reality simulation as described in any of the above embodiments.

[0132] Figure 2 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0133] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0134] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0135] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0136] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB (Universal Serial Bus), network cable, etc.) or wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0137] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0138] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0139] The system described in the above embodiments is used to implement a corresponding augmented reality simulation-based urological surgery training method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0140] Example 4

[0141] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform a urological surgery training method based on augmented reality simulation as described in any of the above embodiments.

[0142] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0143] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute a urological surgery training method based on augmented reality simulation as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0144] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0145] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0146] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0147] Therefore, the units of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0148] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for training urological surgeries based on augmented reality simulation, characterized in that, Includes the following steps: S1. Construct a virtual surgical scene based on a pre-set three-dimensional anatomical model of the urinary system; S2. In the virtual surgical scenario, integrate a real-time interactive digital twin model of surgical instruments; S3. Drive the digital twin model of the surgical instrument to dynamically respond in the virtual surgical scene according to the user's operation input signal; S4. The virtual surgical procedure after dynamic response is superimposed onto the field of view of the real training environment through an augmented reality display device; S5. Based on the user's continuous operation behavior in the overlay field of view, generate corresponding operation trajectory data and store it in the training database.

2. The urological surgery training method based on augmented reality simulation according to claim 1, characterized in that, S2 includes: Obtain key parameters of physical surgical instruments, including: geometric parameters, degrees of freedom of motion, and force feedback characteristics; A dynamic simulation model of the surgical instrument was established based on the aforementioned key parameters; The dynamic simulation model is registered with the spatial coordinate system in the virtual surgical scene to integrate a real-time interactive digital twin model of surgical instruments.

3. The urological surgery training method based on augmented reality simulation according to claim 1, characterized in that, The methods for obtaining user input signals include: The movement and posture data of the trainee's upper limbs and hands are collected using wearable motion capture devices; Based on the motion posture data, the estimated values ​​of operation direction, speed, and contact force are analyzed; The parsing results are converted into operational input signals that can be recognized in the virtual environment.

4. The urological surgery training method based on augmented reality simulation according to claim 3, characterized in that, The steps for driving the digital twin model of the surgical instrument to perform dynamic responses include: The positional changes of the instrument in virtual space are calculated based on the received operational input signals; Determine whether the pose change causes a spatial collision with the virtual anatomical structure; If a collision occurs, the local organ morphology is updated according to the preset tissue deformation model, and the resistance signal is fed back to the user.

5. The urological surgery training method based on augmented reality simulation according to claim 1, characterized in that, S4 includes: The user's viewpoint position and orientation are determined by the spatial positioning module built into the head-mounted display device; Real-time rendering of perspective images of the virtual surgical procedure based on viewpoint information; The perspective image is overlaid as a transparent layer onto the actual operating table view captured by the camera.

6. The urological surgery training method based on augmented reality simulation according to claim 5, characterized in that, The superimposed field of view also includes auxiliary guiding elements, and the method for generating the auxiliary guiding elements includes: Identify key anatomical landmarks based on the current stage of surgery; Retrieve the standard operation path template for this stage from the training database; Based on the deviation between the standard operation path and the user's actual operation trajectory, visual prompts or highlighted areas are generated and superimposed on the field of view.

7. The urological surgery training method based on augmented reality simulation according to claim 1, characterized in that, S5 includes: Real-time recording of the spatial coordinate sequence of the instrument tip during user operation; Simultaneously collect operation time points, estimated applied force values, and angle changes in viewing perspective to generate multi-dimensional time-series data; Multidimensional time-series data is encapsulated into structured operation trajectory data packets using a unified timestamp.

8. A urological surgery training system based on augmented reality simulation, said system being used to implement the method according to any one of claims 1-7, characterized in that, include: Modules for building, integrating, driving, overlaying, and generating; The construction module is used to construct a virtual surgical scene based on a preset three-dimensional anatomical model of the urinary system; The integration module is used to integrate real-time interactive digital twin models of surgical instruments in the virtual surgical scenario; The driving module is used to drive the digital twin model of the surgical instrument to dynamically respond in the virtual surgical scene according to the user's operation input signal; The overlay module is used to overlay the virtual surgical procedure after dynamic response onto the field of view of the real training environment through an augmented reality display device; The generation module is used to generate corresponding operation trajectory data based on the user's continuous operation behavior in the superimposed field of view and store it in the training database.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 7.