Temporal bone side skull base operation simulation method and system and related equipment
By constructing a three-dimensional model of the temporal bone lateral skull base and recording the operation trajectory and mechanical data, and combining deep learning and machine learning technologies, the problem of inaccurate evaluation in existing surgical simulation systems has been solved. This enables accurate quantification and real-time evaluation of the surgeon's operational skills, thereby improving the success rate and safety of the surgery.
Patent Information
- Application Number
- CN202511449423.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-02-10
AI Technical Summary
Existing surgical simulation systems lack accuracy in assessing surgical skill levels, cannot effectively quantify surgeons' skill levels, and suffer from multimodal data silos and computational delays, leading to inaccurate assessment results.
By acquiring medical images of the temporal bone lateral skull base, a three-dimensional model of the target is constructed, and operation trajectory data and mechanical data are recorded. The operation data analysis model is used for feature extraction and fusion, and combined with deep learning and machine learning technologies, surgical operation quality assessment is achieved.
It improves the accuracy of surgical simulation assessment, better quantifies the surgeon's skill level, provides real-time warnings and optimization suggestions, and enhances the success rate and safety of surgery.
Smart Images

Figure CN121506512A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, system and related equipment for simulating surgery on the temporal bone side skull base. Background Technology
[0002] In recent years, the rapid development of medical imaging technology has brought revolutionary breakthroughs to the preoperative planning and simulation training of temporal bone lateral skull base surgery. High-resolution CT can clearly display the temporal bone air cell system and bony landmarks, MRI can accurately identify the fine course of the facial nerve labyrinth segment, and DSA-CT fusion technology has achieved precise registration of blood vessels and bone structures. At the same time, deep learning-based automatic segmentation algorithms have significantly improved the extraction efficiency of submillimeter structures such as the inner ear membranous labyrinth and jugular bulb.
[0003] In related technologies, although surgical simulation systems have achieved a high level of dynamic simulation of tissue structures during surgery, they can only evaluate the effectiveness of surgical simulation based on single-modal data (such as spatial trajectory data) in surgical simulation assessment, resulting in inaccurate assessment results and an inability to effectively and accurately quantify the surgeon's operational skill level.
[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0005] The main objective of this application is to provide a method, system, and related equipment for simulating temporal bone side skull base surgery, which can improve the accuracy of simulated surgical evaluation.
[0006] To achieve the above objectives, one aspect of this application provides a method for simulating temporal bone lateral skull base surgery, the method comprising: Obtain medical images of the lateral skull base of the temporal bone; A model is constructed based on the medical images to obtain a target 3D model; In the simulated surgery on the target three-dimensional model, simulated operation data is recorded, wherein the simulated operation data includes operation trajectory data and mechanical data; The preprocessed operation trajectory data and the mechanical data are input into the trained operation data analysis model for feature extraction to obtain operation trajectory features and mechanical features. The operational trajectory features and the mechanical features are fused to obtain fused features; The quality of surgical procedures was assessed based on the fusion features, and the assessment results were obtained.
[0007] In some embodiments, the medical images include magnetic resonance imaging (MRI) images and computed tomography (CT) images, and the step of model building based on the medical images to obtain a target three-dimensional model includes: The preprocessed MRI images are segmented to obtain the first segmentation data, which represents the bone structure. The preprocessed computed tomography image is segmented into soft tissue structures to obtain second segmentation data; The target 3D model is constructed based on the first segmentation data and the second segmentation data.
[0008] In some embodiments, the simulated operation data is obtained through the following steps: The coordinate positions of the surgical instruments are recorded to obtain multiple three-dimensional coordinates; The three-dimensional coordinates are connected according to the recording time to obtain the operation trajectory data; Multiple mechanical data are recorded using pressure sensors on the target 3D model.
[0009] In some embodiments, the step of inputting the preprocessed operation trajectory data and the mechanical data into a trained operation data analysis model for feature extraction to obtain operation trajectory features and mechanical features includes: Perform small kernel convolution on the operation trajectory data to obtain local trajectory features; Perform large-kernel convolution on the local trajectory features to obtain global trajectory features; The mechanical features are obtained by performing independent convolutions on the mechanical data. The operation trajectory features are obtained by performing residual connection based on the local trajectory features and the global trajectory features.
[0010] In some embodiments, the assessment results include the probability of complication, and the method further includes: The probability of the complication occurrence is input into the complication early warning model to obtain a real-time early warning signal; The step of inputting the probability of complication occurrence into the complication early warning model to obtain a real-time early warning signal includes: The corresponding warning level is determined by matching the probability range corresponding to the preset warning level based on the probability of the complication occurrence. The real-time warning signal is triggered according to the warning level.
[0011] In some embodiments, the operational data analysis model is trained through the following steps: Obtain general surgical datasets and specific surgical datasets; The general surgical dataset is input into the operational data analysis model to obtain the first training result; Based on the first training result, the operational data analysis model is adjusted for the first time to obtain the pre-trained operational data analysis model; The specific surgical dataset is input into the pre-trained operational data analysis model to obtain the second training result; Based on the second training result, the operational data analysis model is adjusted twice to obtain the trained operational data analysis model.
[0012] In some embodiments, the method further includes: The simulated operation data and the evaluation results are analyzed to obtain surgical optimization suggestions.
[0013] To achieve the above objectives, another aspect of this application proposes a surgical simulation method system for the temporal bone lateral skull base, the system comprising: The first acquisition module is used to acquire medical images of the temporal skull base; The modeling module is used to construct a model based on the medical images to obtain a target 3D model. The second acquisition module is used to record simulation operation data during the simulated surgery on the target three-dimensional model, wherein the simulation operation data includes operation trajectory data and mechanical data; The analysis module is used to input the preprocessed operation trajectory data and the mechanical data into the trained operation data analysis model for feature extraction, thereby obtaining operation trajectory features and mechanical features. The fusion module is used to fuse the operation trajectory features and the mechanical features to obtain fused features; The evaluation module is used to evaluate the quality of surgical procedures based on the fusion features and obtain evaluation results.
[0014] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0015] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0016] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0017] The embodiments of this application include at least the following beneficial effects: This application provides a method, system, and related equipment for simulating temporal bone lateral skull base surgery. This method involves acquiring medical images of the temporal bone lateral skull base; constructing a target three-dimensional model based on the medical images; recording simulated operation data during the simulated surgery on the target three-dimensional model, wherein the simulated operation data includes operation trajectory data and mechanical data; inputting the preprocessed operation trajectory data and mechanical data into a trained operation data analysis model for feature extraction to obtain operation trajectory features and mechanical features; fusing the operation trajectory features and mechanical features to obtain fused features; and evaluating the quality of the surgical operation based on the fused features to obtain an evaluation result. By acquiring the operation trajectory data and mechanical data of the simulated surgery on the target three-dimensional model, extracting features from the operation trajectory data and mechanical data, and then fusing the features, this application can better simultaneously reflect the surgeon's trajectory control and force control during surgery, thereby better quantifying the surgeon's operational skill level and improving the accuracy of the evaluation results. Attached Figure Description
[0018] Figure 1 This is a flowchart of the temporal bone lateral skull base surgery simulation method provided in the embodiments of this application; Figure 2 This is an evaluation flowchart of the temporal bone lateral skull base surgery simulation method provided in the embodiments of this application; Figure 3 This is a complete operational schematic diagram of the surgical simulation system provided in the embodiments of this application; Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0020] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0021] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0023] The temporal bone lateral skull base region is one of the most anatomically complex areas in the human body, containing important neurovascular structures such as the inner ear, facial nerve, internal carotid artery, and sigmoid sinus. Due to its deep location, confined space, and proximity to critical functional areas, surgical procedures in this region are extremely challenging. Traditional surgical procedures have a persistently high complication rate, including serious problems such as sensorineural hearing loss, facial paralysis, and cerebrospinal fluid leakage, placing extremely high demands on the surgeon's three-dimensional anatomical knowledge and microsurgical skills. Currently, clinical physician training mainly relies on cadaver specimens; however, the scarcity, high cost, and non-reusability of cadaver specimens severely restrict the widespread adoption of standardized training.
[0024] In recent years, the rapid development of medical imaging technology has brought revolutionary breakthroughs to preoperative planning and simulation training for temporal bone lateral skull base surgery. High-resolution CT can clearly display the temporal bone air cell system and bony landmarks, MRI can accurately identify the subtle course of the facial nerve labyrinth segment, and DSA-CT fusion technology has achieved precise registration of blood vessels and bone structures. At the same time, deep learning-based automatic segmentation algorithms have significantly improved the extraction efficiency of sub-millimeter structures such as the inner ear membranous labyrinth and jugular bulb. However, existing medical imaging systems still suffer from significant multimodal data silos, requiring clinicians to spend a considerable amount of time on manual registration, which not only reduces work efficiency but also increases the risk of human error.
[0025] While significant progress has been made in the application of virtual reality technology in surgical simulation, several key technological bottlenecks remain. Regarding haptic feedback, existing devices lack sufficient force feedback precision, making it difficult to realistically simulate the gradual resistance experienced during bone drilling. In terms of physics engines, computational latency is a common issue when processing high-precision temporal bone models, resulting in a lack of realism in virtual interaction. Related clinical studies have shown that existing virtual reality systems have low mechanical similarity scores to real surgical environments, severely limiting their training effectiveness. Meanwhile, although 3D-printed personalized models can provide physical haptic feedback, they still have significant shortcomings in material properties and functional simulation. Traditional 3D printing materials struggle to accurately simulate the mechanical properties of mastoid processes with varying degrees of vaporization, and cannot achieve real-time deformation feedback during surgery. This prevents surgeons from adjusting surgical strategies based on the model's dynamic response.
[0026] Although surgical simulation systems have achieved a high level of dynamic simulation of tissue structures during surgery, they can only evaluate the effectiveness of surgical simulation based on single-modal data (such as spatial trajectory data) in surgical simulation assessment, resulting in inaccurate assessment results and an inability to effectively and accurately quantify the surgeon's operational skill level.
[0027] In view of this, this application provides a method, system and related equipment for simulating temporal bone lateral skull base surgery. This solution obtains the operation trajectory data and mechanical data of the simulated surgery on the target three-dimensional model, extracts features from the operation trajectory data and mechanical data and then performs feature fusion, which can better reflect the surgeon's trajectory control and force control during the operation, thereby better quantifying the surgeon's operational skill level and improving the accuracy of the evaluation results.
[0028] The temporal bone lateral skull base surgery simulation method provided in this application relates to the field of computer technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited thereto. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application that implements the temporal bone lateral skull base surgery simulation method, but is not limited to the above forms.
[0029] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0030] Figure 1 This is an optional flowchart of the temporal bone lateral skull base surgical simulation method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.
[0031] Step S101: Obtain medical images of the temporal skull base; Step S102: Construct a model based on medical images to obtain a target 3D model; Step S103: During the simulated surgery on the target three-dimensional model, the simulated operation data is recorded, including operation trajectory data and mechanical data. Step S104: Input the preprocessed operation trajectory data and mechanical data into the trained operation data analysis model for feature extraction to obtain operation trajectory features and mechanical features. Step S105: The operation trajectory features and mechanical features are fused to obtain fused features; Step S106: Perform surgical procedure quality assessment based on fusion characteristics to obtain assessment results.
[0032] Steps S101 to S106 as shown in the embodiments of this application, by extracting features from the operation trajectory data and mechanical data and then performing feature fusion, can better reflect the surgeon's trajectory control and force control during the operation, thereby better quantifying the surgeon's operational skill level and improving the accuracy of the evaluation results.
[0033] In step S101 of some embodiments, before the actual surgery, the doctor can perform a simulated surgery on the patient's surgical site, which helps improve the success rate of the surgery. Therefore, it is necessary to first obtain medical images of the patient's temporal skull base, including magnetic resonance imaging (MRI) images and computed tomography (CT) images. For example, in temporal skull base surgery, high-resolution CT scans are used to obtain CT images, thereby acquiring detailed bony structural information of the temporal skull base. 3T MRI multi-sequence scans (T2-SPACE, CISS, etc.) are used to obtain MRI images, thereby acquiring structural information of soft tissues, nerves, blood vessels, etc. Obtaining different tissue structures of the human body through different images provides high-precision information for subsequent three-dimensional modeling.
[0034] In some embodiments, step S102 may include, but is not limited to, steps S201 to S203: Step S201: Segment the preprocessed MRI image into bony structures to obtain the first segmentation data; Step S202: Segment the soft tissue structure of the preprocessed computed tomography image to obtain the second segmentation data; Step S203: Construct the target 3D model based on the first segmentation data and the second segmentation data.
[0035] In steps S201 to S203 of some embodiments, CT and MRI images undergo preprocessing such as denoising, registration, and fusion to improve the spatial consistency of multimodal images. In temporal skull base surgery, before image segmentation, an intelligent segmentation model targeting the anatomical structures of the temporal skull base is constructed based on deep learning network architectures such as U-Net and V-Net. By collecting a large amount of labeled CT and MRI image data of the temporal skull base and setting key anatomical landmarks such as the facial nerve, internal carotid artery, and inner ear structures (cochlea, vestibule, semicircular canals, etc.) for focused segmentation, the intelligent segmentation model is trained and optimized. The segmentation accuracy of the intelligent segmentation model is evaluated using metrics such as the Dice coefficient and Hausdorff distance, achieving millimeter-level accuracy. The trained intelligent segmentation model segments the bony structures in the MRI images to obtain first segmentation data, which can represent information about different bone structures. The intelligent segmentation model then segments the soft tissue structures in the computed tomography images to obtain second segmentation data, which can represent structural information about soft tissues, nerves, and blood vessels in different locations. Based on the first and second segmentation data, a target 3D model of the temporal bone lateral skull base is constructed using surface rendering or volume rendering techniques. The different types of segmentation data allow the target 3D model to simulate the pulsation of blood vessels such as the internal carotid artery and the conduction process of nerves such as the facial nerve, supporting visualization. It also supports interactive operations such as rotation, scaling, and translation, facilitating user observation and analysis. This embodiment leverages the complementary advantages of CT and MRI images to improve the accuracy of the 3D model. Through intelligent segmentation, it achieves automatic identification and segmentation of key anatomical structures, improving segmentation efficiency and accuracy, reducing manual intervention, and constructing a realistic and interactive 3D model, providing a reliable anatomical basis for temporal bone lateral skull base surgical simulation.
[0036] In some embodiments, the target 3D model includes a 3D virtual model and a 3D solid model. Based on the preceding medical images, a 3D virtual model is constructed. Then, based on the 3D virtual model, the solid model printing path is planned and the support structure is designed. A multi-nozzle bioprinter is used to print the 3D solid model. For the bony structures of the 3D solid model, a high-hardness light-cured resin is used to simulate the hardness and strength of the temporal bone. For the soft tissues of the 3D solid model, hydrogels with different elastic moduli are used to simulate the texture of soft tissues such as muscles, nerves, and blood vessels. By precisely controlling the ratio of different materials and printing parameters, a gradient transition between bony structures and soft tissues is achieved, simulating real anatomical structures. Miniaturized, high-sensitivity pressure sensors, such as piezoresistive or capacitive sensors, are selected. During the printing process, the pressure sensor array is embedded at the interface between soft tissue and bony structures, as well as around key anatomical structures. In simulated surgery, the pressure sensor array acquires pressure sensor signals in real time, performs data processing and analysis, visualizes the operating force, and provides real-time feedback to help doctors control the operating force and avoid damage to important structures.
[0037] In step S103 of some embodiments, during the simulated surgery on the target 3D model, the user enters a highly immersive and realistically interactive digital twin platform using a VR headset (such as Meta Quest Pro and HTC Vive). This platform provides doctors with a space for preoperative planning, simulation training, and surgical rehearsals. The user uses VR controllers or force feedback devices to grasp, move, and rotate virtual surgical instruments and interact with virtual tissues. Based on game engines such as Unity and Unreal Engine, combined with physics engines (such as Havok and PhysX), the physical properties of tissues are simulated, including bone hardness, soft tissue elasticity, and blood vessel fragility. High-precision haptic feedback devices (such as Geomagic Touch and HaptX Gloves) are used to simulate the feel of surgical instruments, providing force feedback and simulating the realistic tactile sensations of bone grinding, cutting, and traction. In simulating dynamic bleeding effects, a real-time fluid dynamics algorithm is developed based on the Navier-Stokes equations to simulate blood flow and the bleeding process. Based on a vascular network model, which defines parameters such as vascular pressure and flow, the system calculates bleeding rate and volume in real time according to surgical procedures and displays the results visually. When users perform hemostasis procedures using virtual instruments, such as electrocoagulation and compression, the system provides real-time feedback on the hemostasis effect.
[0038] In some embodiments, the simulated operation data can be obtained through steps S301 to S303.
[0039] Step S301: Record the coordinate position of the surgical instruments to obtain multiple three-dimensional coordinates; Step S302: Connect the three-dimensional coordinates according to the recording time to obtain the operation trajectory data; Step S303: Record multiple mechanical data using pressure sensors on the target 3D model.
[0040] In steps S301 to S303 of some embodiments, the target 3D model includes a 3D virtual model and a 3D solid model. The 3D virtual model and the 3D solid model are spatially aligned in a VR headset. During the user's surgical operation, the 3D virtual model and the 3D solid model change synchronously. The spatial pose of surgical instruments (such as drills and peelers) is captured in real time using a six-degree-of-freedom electromagnetic tracking system to obtain the 3D coordinates (x, y, z) and Euler angles (α, β, γ) of the instrument tip. The Euler angles represent the instrument orientation. Based on timestamps, the discrete 3D coordinates are connected to generate a continuous trajectory curve (operation trajectory data), which also includes Euler angle data. Pressure data is recorded using a pressure sensor array of the 3D solid model.
[0041] In some embodiments, resistance data collected by a micro-pressure sensor array in the 3D solid model is fed back to the PhysX physics engine via a dedicated interface, dynamically correcting the resistance gradient curve of the virtual tissue deformation algorithm. Based on the resistance data and the vascular bleeding parameters calculated using the Navier-Stokes equations, the fluid flow rate of the microfluidic system embedded in the 3D virtual model is driven in real time, achieving biomimetic simulation of the actual bleeding scenario. This bidirectional transmission forms a closed-loop optimization mechanism, significantly improving the simulation fidelity of the system. Dynamic process simulation is dominated by the virtual platform, which, based on real-time fluid dynamics algorithms, can simulate complex intraoperative scenarios (such as blood flow diffusion in a ruptured sigmoid sinus), while the 3D solid model, limited by its static material properties, cannot achieve such dynamic responses. Biomechanical feedback is provided by the 3D solid model; the hydrogel material recreates the tactile sensation of nerves and blood vessels, effectively compensating for the difference between the virtual force feedback device and the tactile sensation of real tissue. The two work together to construct a comprehensive training environment.
[0042] In some embodiments, step S104 is described in detail below. Figure 2 After acquiring the spatial trajectory of surgical instruments (operation trajectory data), contact pressure (mechanical data), and bone conduction spectrum data (acoustic data), these heterogeneous data are input into the preprocessing module for standardization. For the operation trajectory data, adaptive Kalman filtering is used to eliminate physiological hand tremors while preserving the operator's intention signal. For the mechanical-acoustic data, dynamic time warping is used to align multi-source timing sequences and compensate for sensor delay differences. Combining operation trajectory data, mechanical data, and acoustic data to comprehensively assess the surgeon's skill level helps improve the objectivity and accuracy of the assessment.
[0043] In some embodiments, step S104 may include, but is not limited to, steps S401 to S404: Step S401: Perform small kernel convolution on the operation trajectory data to obtain local trajectory features; Step S402: Perform large kernel convolution on the local trajectory features to obtain the global trajectory features; Step S403: Perform independent convolutions on the mechanical data to obtain mechanical features; Step S404: Perform residual connection based on local trajectory features and global trajectory features to obtain operation trajectory features.
[0044] In steps S401 to S404 of some embodiments, the operational data analysis model ResNet-3D is improved with temporal bone specificity. This architecture is designed with two parallel analysis pathways, including a spatiotemporal feature extraction branch and a mechanical analysis branch. The spatiotemporal feature extraction branch includes large-kernel convolutional layers and small-kernel convolutional layers. Operational trajectory data first enters the small-kernel convolutional layer, and an improved 3×3×3 small-kernel convolution is used to capture the microscopic operational details of the operational trajectory data to obtain local trajectory features. The local trajectory features are then input into the large-kernel convolutional layer, and a deep 7×7×5 large convolution identifies the overall motion pattern to obtain global trajectory features. The local trajectory features and global trajectory features are input into residual blocks for residual concatenation. Attention gating mechanisms between residual blocks can actively strengthen the feature weights of high-risk anatomical regions (such as the ampulla of the semicircular canal). Independent convolution of mechanical data can reduce the interference of spatial information on mechanical feature extraction.
[0045] In step S105 of some embodiments, the operation trajectory features and mechanical features are fused across modally in the residual layer of the network. A spatial-mechanical coupling tensor is generated by graph convolution. Through physical-spatial essential correlation modeling and dynamic weight allocation, cross-modal feature collaboration in the surgical scenario is realized. This can better reflect the surgeon's trajectory control and force control during surgery, thereby better quantifying the surgeon's operational skill level. In addition, this embodiment can also use acoustic data as a reference factor for operational quality assessment.
[0046] In some embodiments, the operational data analysis model can be trained through steps S410 to S450.
[0047] Step S410: Obtain the general surgical dataset and the specific surgical dataset; Step S420: Input the general surgical dataset into the operation data analysis model to obtain the first training result; Step S430: Based on the first training result, the parameters of the operational data analysis model are adjusted for the first time to obtain the preliminary trained operational data analysis model; Step S440: Input the specific surgical dataset into the pre-trained operational data analysis model to obtain the second training result; Step S450: Adjust the parameters of the operational data analysis model a second time based on the second training result to obtain the trained operational data analysis model.
[0048] In steps S410 to S450 of some embodiments, a large amount of surgical operation trajectory data from expert physicians is collected for training and optimizing the operation data analysis model. A two-stage transfer learning strategy is employed for training the operation data analysis model. First, the basic feature extraction capability is pre-trained on the publicly available JIGSAWS general surgical dataset. The first training result includes the surgical score and the probability of complication. Based on the output of the first training result and the actual score and complication label of the general surgical dataset, a first loss function is constructed to obtain the first loss value. By adjusting the parameters of the operation data analysis model, the first loss value is reduced, resulting in a pre-trained operation data analysis model. Then, fine-tuning is performed based on 500 cases of temporal bone surgery data (a specific surgical dataset) (containing 28 anatomical variations). The second training result also includes the surgical score and the probability of complication. A second loss function is constructed, combining mean squared error (score) and focal loss (complication prediction), and adding an anatomical constraint term (penalizing instrument intrusion into high-risk areas). The parameters of the operation data analysis model are adjusted again to minimize the second loss value, resulting in a well-trained operation data analysis model. The ResNet-3D network serves as the core technology carrier for operational data analysis models, with its improved design specifically tailored to the precision characteristics of temporal bone surgery. This architecture captures micrometer-level motion features through multi-scale spatiotemporal convolutional modules, enhances the perception of high-risk areas such as the facial nerve canal through attention mechanisms between residual blocks, and simultaneously achieves operational quality assessment and risk prediction through multi-task output layers. The TensorRT lightweight model is deployed and optimized based on the ResNet-3D network, transforming the complex computational process of the original network into an inference engine that meets the real-time requirements of surgery through graph structure compression and mixed-precision quantization.
[0049] In step S106 of some embodiments, please refer to Figure 2The evaluation results include an operational score and the probability of complication. The fused features are input into the operational quality assessment layer, and these features are analyzed and evaluated based on multiple indicators, such as operational fluency (continuity index based on trajectory acceleration integral), instrument stability (jitter energy spectral density in the 0.5-3Hz frequency band), anatomical positioning accuracy (Hausdorff distance between the instrument tip and the marked point), and risk contact frequency (number of unauthorized approaches to high-risk areas (facial nerve) per unit time, and abnormal bleeding volume due to pressure values). The operational score is generated as a percentage score by combining the preceding indicators, with the weight parameters of different indicators dynamically adjusted according to the anatomical difficulty (e.g., increased risk control weight for high jugular bulb cases). The probability of complication is output through a Sigmoid activation function based on the fused features. The operational score can comprehensively quantify the surgeon's skill level, providing an intuitive reference, while the probability of complication can serve as a trigger indicator for subsequent early warnings, enabling real-time feedback during surgery.
[0050] In some embodiments, the temporal skull base surgical simulation method further includes step S107: Step S107: Input the probability of complication occurrence into the complication early warning model to obtain a real-time early warning signal.
[0051] In some embodiments, step S107 may include, but is not limited to, steps S701 to S702: Step S701: Match the probability range corresponding to the preset warning level with the probability of complication occurrence to determine the corresponding warning level; Step S702: Trigger a real-time warning signal based on the warning level.
[0052] In steps S701 to S702 of some embodiments, the complication warning model presets multiple real-time warning signals. Each real-time warning signal can represent different levels of operational warning and corresponds to different complication probability ranges. When a dangerous operation mode is identified, a warning prompt is issued in a timely manner, and corresponding avoidance suggestions are given. This embodiment uses a multimodal feedback mechanism to realize the warning prompt. Taking the accidental opening of the semicircular canal as an example, when the drilling depth is detected to be close to the critical value of the bone wall, the probability of complication is greater than 90%. The corresponding area in the surgical field will immediately show a red flashing mark, and the drill resistance will simulate the real penetration feeling, accompanied by a warning sound of a specific frequency. Higher-level warnings will trigger active intervention, such as automatically reducing the drilling speed or recalling the repair plan guidance. In response to the risk of facial nerve thermal injury, the system not only displays the outline of the nerve course, but also simulates the cooling process through virtual saline spray. All warning information has been optimized by human factors to effectively transmit risk information without interfering with the main surgical field.
[0053] In some embodiments, the temporal skull base surgical simulation method further includes step S108: Step S108: Analyze the simulation operation data and evaluation results to obtain surgical optimization suggestions.
[0054] In step S108 of some embodiments, big data analysis is performed on the surgeon's operation, and the simulated operation data includes operation trajectory, instrument use, and occurrence of complications. Based on machine learning and big data analysis technology, the surgeon's operation data is analyzed to uncover operation patterns, identify risk factors, and evaluate surgical outcomes. Based on the surgeon's operation data, personalized training suggestions, surgical plan optimization, and complication prevention measures are provided.
[0055] In some embodiments, during facial nerve decompression via the middle cranial fossa, the surgeon needs to open the facial nerve canal under a microscope, with an operating space of only 2-3 mm, requiring instrument movement precision of 0.1 mm. A resident physician attempted this procedure for the first time. The system uses a 9-axis IMU sensor (sampling rate 1 kHz) integrated on the microsurgical instruments to acquire operational data in real time: establishing a three-dimensional coordinate system with the geniculate ganglion as the origin, recording the coordinates of the dissecting tip, and calculating velocity (0-2 mm / s) and acceleration (±0.5 m / s²). 2 The ResNet-3D model monitored contact force (0.01-0.5N) and torque (0-15N·mm) and motion smoothness index. The ResNet-3D model identified the following features: abnormal trajectory fluctuations (frequency 8-12Hz, amplitude 0.3mm) at a certain moment, the distance between the instrument and the lower wall of the facial nerve canal shortened to 0.2mm (safe threshold 0.5mm), and the contact pressure gradient suddenly increased (ΔF / Δt=0.25N / s).
[0056] Multi-level early warning system response: Primary warning (risk probability 65%): The outline of the facial nerve canal turns amber in the microscope field of view; tactile feedback generates a 0.1N resistance to simulate touch. Advanced warning (risk probability 89%): When a pushing motion lasting 3 seconds is detected (pressure > 0.3 N and displacement < 0.1 mm); AR projection displays the nerve fiber pathways (red arrows); the drill automatically enters safety mode (rotation speed reduced to 5000 rpm). Finally, the system calculates a real-time score based on 12 parameters, as shown in Table 1: Table 1 Scoring Sheet
[0057] The postoperative report showed that the operation score was lower than the expert benchmark (85 points), with the main reasons for the low score being the accuracy of anatomical identification and the fluency of the operation.
[0058] In some embodiments, when dealing with the sigmoid sinus surface bone during radical mastoidectomy, a balance must be struck between bone removal efficiency and venous sinus protection. An attending physician trains on a simulator to perform the procedure under a variant anatomy (anterior sinus view).
[0059] The system uses multimodal monitoring to record the electric drill's trajectory, detect drilling resistance (50-200N) and temperature changes, as well as bone conduction spectrum characteristics.
[0060] The operational data analysis model identified that at a certain moment, there was a jump feed (displacement of 0.5 mm / time, interval of 200 ms), the angle between the drill bit and the tangent of the sinus surface fluctuated (35°±15°), and the acoustic spectrum energy suddenly increased by 6 dB at 3.8 kHz (indicating bone penetration).
[0061] Graded response mechanism: 1) Spatial warning: When the drill bit is <0.3mm from the sinus wall, the microscope field of view shows a pulsating red halo; 2) Mechanical warning: When the resistance drops by 30%, the tactile feedback produces a "feeling of emptiness"; the virtual hemostatic gauze is automatically activated (pressure maintained at 20mmHg); 3) Comprehensive warning: Three-dimensional holographic projection displays the simulation of sinus wall damage.
[0062] The system generated an operation quality report: 1) Bone removal uniformity: 85% (target >90%); 2) Sinus wall exposure area: 2.3mm 2 (Safety range <3mm) 2 ); 3) Number of dangerous operations: 1 time (drill bit slippage); 4) Temperature control: ΔT=2.7℃ (safety threshold<5℃).
[0063] Skills matrix assessment: 1) Handling anatomical variations: Excellent (88th percentile); 2) Drilling rhythm control: Good (83rd percentile); 3) Bleeding risk management: Superior (94th percentile); 4) Instrument switching efficiency: Meets the standard (79th percentile).
[0064] The surgeon's final score was 83 points. The system recommended strengthening the "thin-layer continuous drilling of bone cortex" special training (training module #7) and reminded the operator to pay attention to maintaining the drill bit angle (recommended 45±5°).
[0065] In some embodiments, please refer to Figure 3A surgical simulation system includes a multimodal image modeling module, a digital twin interactive platform, a 3D bioprinting system, an intelligent evaluation system, and a cloud platform collaboration system. The multimodal image modeling module is a key component in constructing the temporal bone lateral skull base surgery simulation system. Its core technologies include multimodal image fusion, intelligent segmentation algorithms, and 3D dynamic modeling. The development of this module will provide important support for the precision and minimally invasive nature of temporal bone lateral skull base surgery. The digital twin interactive platform is the core module of the temporal bone lateral skull base surgery simulation system. Its core technologies include a physics engine, haptic feedback, and fluid dynamics algorithms. The development of this platform will also provide important support for the precision and minimally invasive nature of temporal bone lateral skull base surgery. The 3D bioprinting training system uses multi-material gradient printing technology (photocurable resin + hydrogel) to achieve biomimetic reproduction of bony structures and soft tissues, and embeds a micro pressure sensor array to monitor the operational force in real time. The virtual surgical scene in the digital twin interactive platform and the 3D bioprinted physical model share the same patient-specific anatomical data source. Both use unified quantitative standards to achieve consistency in the spatial attributes and mechanical properties between virtual pre-simulation and physical operation. The intelligent assessment system is a crucial component of the temporal bone lateral skull base surgery simulation system. Its core technologies include deep learning, machine learning, and quantitative assessment, outputting evaluation results of the surgeon's actions. Based on these evaluation results, the cloud platform collaboration system enables big data analysis of surgeon actions, collaborative planning of multi-center surgical procedures, and VR remote consultations, promoting communication, learning, and progress in temporal bone lateral skull base surgery. By constructing a closed-loop training system encompassing "image modeling - virtual pre-performance - physical operation - digital debriefing," the safety of lateral skull base surgery is significantly improved.
[0066] This application also provides a surgical simulation method system for temporal bone lateral skull base surgery, which can implement the above-described method. The system includes: The first acquisition module is used to acquire medical images of the temporal skull base; The modeling module is used to build models based on medical images to obtain the target 3D model; The second acquisition module is used to record simulation operation data during the simulated surgery on the target three-dimensional model, wherein the simulation operation data includes operation trajectory data and mechanical data. The analysis module is used to input the preprocessed operation trajectory data and mechanical data into the trained operation data analysis model for feature extraction, thereby obtaining operation trajectory features and mechanical features. The fusion module is used to fuse operational trajectory features and mechanical features to obtain fused features. The evaluation module is used to assess the quality of surgical procedures based on fusion features and obtain evaluation results. It is understood that the content of the above method embodiments is applicable to this system embodiment, and the specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0067] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0068] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0069] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the methods described in the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0070] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0071] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0072] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0073] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0074] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0075] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0076] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0077] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0078] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0079] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0080] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0081] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0082] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0083] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0084] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0085] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for simulating surgery on the temporal bone lateral skull base, characterized in that, The method includes the following steps: Obtain medical images of the lateral skull base of the temporal bone; A model is constructed based on the medical images to obtain a target 3D model; In the simulated surgery on the target three-dimensional model, simulated operation data is recorded, wherein the simulated operation data includes operation trajectory data and mechanical data; The preprocessed operation trajectory data and the mechanical data are input into the trained operation data analysis model for feature extraction to obtain operation trajectory features and mechanical features. The operational trajectory features and the mechanical features are fused to obtain fused features; The quality of surgical procedures was assessed based on the fusion features, and the assessment results were obtained.
2. The method according to claim 1, characterized in that, The medical images include MRI images and computed tomography (CT) images. The step of constructing a model based on the medical images to obtain a target 3D model includes: The preprocessed MRI images are segmented to obtain the first segmentation data, which represents the bone structure. The preprocessed computed tomography image is segmented into soft tissue structures to obtain second segmentation data; The target 3D model is constructed based on the first segmentation data and the second segmentation data.
3. The method according to claim 1, characterized in that, The simulated operation data is obtained through the following steps: The coordinate positions of the surgical instruments are recorded to obtain multiple three-dimensional coordinates; The three-dimensional coordinates are connected according to the recording time to obtain the operation trajectory data; Multiple mechanical data are recorded using pressure sensors on the target 3D model.
4. The method according to claim 1, characterized in that, The preprocessed operation trajectory data and mechanical data are input into a trained operation data analysis model for feature extraction to obtain operation trajectory features and mechanical features, including: Perform small kernel convolution on the operation trajectory data to obtain local trajectory features; Perform large-kernel convolution on the local trajectory features to obtain global trajectory features; The mechanical features are obtained by performing independent convolutions on the mechanical data. The operation trajectory features are obtained by performing residual connection based on the local trajectory features and the global trajectory features.
5. The method according to claim 1, characterized in that, The assessment results include the probability of complication, and the method further includes: The probability of the complication occurrence is input into the complication early warning model to obtain a real-time early warning signal; The step of inputting the probability of complication occurrence into the complication early warning model to obtain a real-time early warning signal includes: The corresponding warning level is determined by matching the probability range corresponding to the preset warning level based on the probability of the complication occurrence. The real-time warning signal is triggered according to the warning level.
6. The method according to claim 1, characterized in that, The operational data analysis model is trained through the following steps: Obtain general surgical datasets and specific surgical datasets; The general surgical dataset is input into the operational data analysis model to obtain the first training result; Based on the first training result, the operational data analysis model is adjusted for the first time to obtain the pre-trained operational data analysis model; The specific surgical dataset is input into the pre-trained operational data analysis model to obtain the second training result; Based on the second training result, the operational data analysis model is adjusted twice to obtain the trained operational data analysis model.
7. The method according to claim 1, characterized in that, The method further includes: The simulated operation data and the evaluation results are analyzed to obtain surgical optimization suggestions.
8. A surgical simulation system for the temporal bone lateral skull base, characterized in that, The system includes: The first acquisition module is used to acquire medical images of the temporal skull base; The modeling module is used to construct a model based on the medical images to obtain a target 3D model. The second acquisition module is used to record simulation operation data during the simulated surgery on the target three-dimensional model, wherein the simulation operation data includes operation trajectory data and mechanical data; The analysis module is used to input the preprocessed operation trajectory data and the mechanical data into the trained operation data analysis model for feature extraction, thereby obtaining operation trajectory features and mechanical features. The fusion module is used to fuse the operation trajectory features and the mechanical features to obtain fused features; The evaluation module is used to evaluate the quality of surgical procedures based on the fusion features and obtain evaluation results.
9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.