A method, device and storage medium for virtual endoscope digital twin
By constructing a patient-specific three-dimensional anatomical model and extracting semantic features in real time, the problem of blind perspective in soft tissue surgery using virtual laparoscopy technology has been solved. Automatic linkage and path advancement of the virtual laparoscopic perspective have been achieved, improving the reliability and intelligence of surgical navigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING FUDIMAI DIGITAL TECH CO LTD
- Filing Date
- 2026-03-16
- Publication Date
- 2026-05-26
Smart Images

Figure CN121837557B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing and surgical navigation technology, and in particular to a virtual laparoscopic digital twin method, device and storage medium. Background Technology
[0002] Laparoscopic minimally invasive surgery has been widely implemented in general surgery, thoracic surgery, and urology. Its core relies on real-time intraoperative observation of two-dimensional laparoscopic video, requiring surgeons to mentally construct three-dimensional spatial relationships based on experience to make surgical judgments and plan procedures. With the development of computer vision and 3D reconstruction technology, constructing patient-specific 3D models based on preoperative medical images (such as CT and MRI) and using them for surgical planning and intraoperative navigation has become an important direction for improving surgical precision.
[0003] Currently, intraoperative navigation technologies, represented by augmented reality (AR), primarily use optical or electromagnetic tracking to rigidly spatially register preoperative 3D models with the patient's intraoperative position and overlay the virtual model onto the surgical field, providing surgeons with "X-ray vision"-like visual assistance. For example, in spinal or neurosurgical procedures, such systems can significantly improve the precision of procedures such as screw implantation.
[0004] However, the core logic of this technology is to pursue precise alignment between the virtual model and the physical world in geometric space. Its effectiveness highly depends on stable tracking and registration accuracy. In soft tissue surgical scenarios such as abdominal and pelvic surgeries, organs undergo significant deformation and displacement due to respiration, heartbeat, and surgical operations, leading to rigid registration failure and deviations between the virtual model and the real anatomical structure—the so-called "dynamic compensation" problem. Although some studies have attempted to predict and compensate using intraoperative imaging or mechanical data, the problem has not yet been fundamentally solved.
[0005] On the other hand, digital twin technology has brought a new paradigm of full-process simulation to surgery. By constructing a "digital twin" model consistent with the patient's anatomy, it can be used not only for visualization but also for surgical simulation and risk prediction. In existing technologies, the application of digital twins is mostly concentrated in preoperative planning (such as lesion segmentation, vascular reconstruction, and optimal path calculation) and postoperative analysis, or for training robot control models. In the intraoperative phase, although some studies have explored "decomposing" the surgical scene to assist navigation matching and using digital twin models for real-time calculations such as trocar placement, these methods still focus on static or parametric adjustments based on geometric or physical states, rather than a high-level semantic understanding of the surgical process itself.
[0006] Specifically, in the field of virtual laparoscopy (or virtual endoscopy) technology, traditional applications often involve generating a roaming view within a cavity from three-dimensional volumetric data for anatomical teaching or simulation training. The path is typically manually set or automatically generated based on geometric rules such as the centerline, primarily focusing on path smoothness and geometric safety constraints such as "preventing wall perforation." However, these methods lack semantic correspondence with actual surgical steps, instrument interactions, and tissue deformation states. In complex laparoscopic surgeries, such as radical resection of gastric or rectal cancer, the surgery needs to proceed in stages within narrow anatomical spaces according to a specific sequence (such as vascular ligation, fascia dissection, and space separation). Existing virtual laparoscopic systems cannot understand these clinical semantic stages; therefore, their viewpoint movement is blind and detached, unable to accurately reproduce or predict the observation path of the endoscope during real surgery, and difficult to achieve stable linkage with the actual surgical procedure.
[0007] For example, CN103356155B proposes cavity modeling and lesion localization based on virtual endoscopy, which has path browsing capabilities. However, it is mainly based on visual inspection and lacks high-level semantic constraints and perspective control around the process, and its adaptability to tissue deformation and situational changes is limited. CN107248191A provides an automatic / interactive path planning process for complex cavities, covering two-dimensional segmentation, three-dimensional modeling, central path planning and roaming safety control. It focuses on geometric paths and "anti-wall breaking" constraints, but it is insufficient in depicting the dynamic linkage of instrument-tissue interaction, stage transition and semantic driving, and it is difficult to stably correspond to the real surgical procedure. CN101856264B proposes an orthopedic surgical navigation template based on three-dimensional reconstruction, which achieves bone surface guidance through splint and probe structures. It is more inclined to rigid structures and mechanical navigation, and does not cover soft tissue scenes, virtual lens control and semantic navigation. Summary of the Invention
[0008] To address the shortcomings of existing technologies, this invention provides a virtual laparoscopic digital twin method, device, and storage medium. It solves the problems in existing technologies that only provide two-dimensional laparoscopic videos or static three-dimensional models, lack explicit modeling of the semantic state of the surgical scene in the time series, and lack a control mechanism to link the semantic state with the virtual laparoscopic path parameters, resulting in the virtual perspective not being able to advance automatically and stably with the progress of the surgery.
[0009] According to an embodiment of the present invention, a method for creating a virtual laparoscopic digital twin is provided, the method comprising the following steps:
[0010] Based on the acquired preoperative medical imaging data of the patient, a three-dimensional anatomical model of the patient is constructed; in the coordinate system of the three-dimensional anatomical model, a spatial travel path of the virtual endoscope is constructed, and the spatial travel path is divided into several path sub-intervals associated with each expected clinical scenario;
[0011] At least one key node is selected within each of the path sub-intervals, and a set of anatomical structures expected to be visible for each key node is predefined; the spatial travel path, the path sub-intervals, the key nodes, and the set of anatomical structures are integrated to form a path prior model;
[0012] During the surgery, the video sequence of the virtual endoscope acquired in real time is segmented to obtain segmentation results including anatomical structures and surgical instruments; based on the segmentation results of multiple consecutive frames within a preset time window, semantic features reflecting the dynamic changes of the surgical scene are extracted, and the current state of the surgical scene is determined according to the semantic features.
[0013] Based on the current scene state, the corresponding path sub-interval is determined from the path prior model, and the propulsion parameters of the virtual cavity mirror on the spatial travel path are updated accordingly. When the propulsion parameters approach a certain key node predefined in the path prior model, a viewpoint switch is triggered, and the viewpoint of the virtual cavity mirror is switched to the pose corresponding to the key node.
[0014] On the other hand, according to embodiments of the present invention, a virtual endoscopic digital twin device is also provided, employing the aforementioned virtual endoscopic digital twin method. The device includes a data input interface, an image acquisition unit, and a processor. The processor includes a 3D reconstruction and path encoding unit, an instance segmentation and semantic feature extraction unit, a scene determination unit, and a path advancement and view control unit, wherein:
[0015] The data input interface is used to acquire the patient's preoperative medical imaging data and send the imaging data to the three-dimensional reconstruction and path coding module;
[0016] The image acquisition device is used to acquire video sequences of the virtual cavity mirror in real time and send the video sequences to the instance segmentation and semantic feature extraction module;
[0017] The three-dimensional reconstruction and path coding unit is used to construct a three-dimensional anatomical model of the patient based on the received image data; it is also used to construct the spatial travel path of the virtual endoscope, divide the path into sub-intervals and set several key nodes in the coordinate system of the three-dimensional anatomical model to form a path prior model.
[0018] The instance segmentation and semantic feature extraction unit is used to perform instance segmentation on the received video sequence to obtain segmentation results including anatomical structures and surgical instruments; it is also used to extract semantic features reflecting the dynamic changes of the surgical scene based on the segmentation results of multiple consecutive frames within a preset time window.
[0019] The scene determination unit is used to determine the current scene state of the surgery based on the semantic features sent by the instance segmentation and semantic feature extraction module.
[0020] The path advancement and view control unit is used to determine the corresponding path sub-interval from the path prior model based on the current scene state sent by the scene determination module, and update the advancement parameters of the virtual cavity mirror on the spatial travel path accordingly; it is also used to trigger view switching when the advancement parameters approach a certain key node predefined in the path prior model, and switch the view of the virtual cavity mirror to the pose corresponding to the key node.
[0021] In another aspect, according to embodiments of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the aforementioned virtual endoscopic digital twin method.
[0022] Compared with the prior art, the present invention has the following beneficial effects:
[0023] (1) At the path level, by quantifying the laparoscopic route with stable geometric regularity in clinical practice into a generalizable spatial curve and key node set based on the patient-specific three-dimensional anatomical model, the limitations of traditional subjective planning based on surgeon experience or two-dimensional images have been changed, providing a standardized "digital roadmap" for the advancement of virtual perspective, which significantly improves the repeatability and individual adaptability of surgical planning.
[0024] (2) At the semantic level, by performing instance segmentation on real-time endoscopic video and extracting dynamic semantic features based on the segmentation results of multiple consecutive frames within the time window, the complex and ever-changing soft tissue surgical scene is mapped into a finite and clear set of semantic states. This overcomes the dependence of traditional methods on static images or rigid registration, and can more robustly identify the current stage of the surgery, providing reliable perceptual input for subsequent perspective linkage.
[0025] (3) At the control level, through the chain mapping of "semantic state determination → path sub-interval matching → advancement parameter update → key node perspective switching", the automatic linkage of scene-path-perspective is established, so that the virtual endoscopic trajectory can approximately reproduce the observation logic of real surgery on the time axis and realize the "synchronous advancement" of perspective with the surgical process, rather than passively following or manually adjusting, which greatly improves the smoothness and intelligence of navigation.
[0026] (4) At the engineering implementation level, the virtual endoscope perspective switching is driven by video semantic recognition, which has low dependence on additional hardware and can be operated with only weak registration or conventional endoscope equipment, significantly reducing the deployment cost and clinical use threshold of the system.
[0027] (5) At the engineering promotion level, through modular path prior model and interface design, it is easy to extend to different surgical procedures and is applicable to various scenarios such as surgical teaching, skills training and postoperative review, with good scalability and maintainability. Attached Figure Description
[0028] Figure 1 This is a flowchart illustrating the steps of a virtual endoscopic digital twin method according to an embodiment of the present invention.
[0029] Figure 2 This is a schematic diagram of the preoperative three-dimensional model and virtual endoscopic path prior and key nodes in an embodiment of the present invention.
[0030] Figure 3 This is a schematic diagram of the semantic encoding of virtual endoscope path nodes and the expected field of view in an embodiment of the present invention.
[0031] Figure 4 This is a schematic diagram illustrating intraoperative instance segmentation, temporal window semantic feature extraction, and scene determination in an embodiment of the present invention.
[0032] Figure 5 This is a schematic diagram of the overall architecture of a virtual endoscopic digital twin device according to another embodiment of the present invention.
[0033] Figure 6 This is a schematic diagram of the hardware structure of a virtual endoscopic digital twin device according to another embodiment of the present invention.
[0034] In the above figures: 1. Data input interface; 2. Image acquisition device; 3. Processor; 4. Memory; 5. Display terminal; 6. System bus; 7. Image processor; 8. Network interface module; 31. 3D reconstruction and path encoding unit; 32. Instance segmentation and semantic feature extraction unit; 33. Scene determination unit; 34. Path advancement and view control unit; 35. Human-computer interaction and log unit. Detailed Implementation
[0035] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0036] like Figures 1-4 As shown in the figure, an embodiment of the present invention proposes a virtual endoscope digital twin method, the method comprising the following steps:
[0037] S1. Based on the acquired preoperative medical imaging data of the patient, construct a three-dimensional anatomical model of the patient; under the coordinate system of the three-dimensional anatomical model, construct a spatial travel path of the virtual endoscope, and divide the spatial travel path into several path sub-intervals associated with each expected clinical scenario;
[0038] The construction of a three-dimensional anatomical model of the patient based on preoperative medical imaging data specifically includes:
[0039] The patient's preoperative medical imaging data is acquired, and the medical imaging data is segmented into multiple anatomical structures to obtain a three-dimensional model including several anatomical labels. The medical imaging data includes three-dimensional CT data.
[0040] Specifically, based on the reconstruction of publicly available computed tomography (CT) volume data, the patient's preoperative three-dimensional CT data is obtained. The three-dimensional CT data is then segmented into multiple anatomical structures, i.e., the volume data is processed. Through the segmentation function Mapped to multiple anatomical tags This allows us to obtain a three-dimensional model containing multiple anatomical labels. This is used for subsequent construction of virtual endoscope paths and field of view simulation. Among them, This may include the sigmoid mesentery, rectal mesentery, inferior mesenteric artery / vein (IMA / IMV, hereinafter collectively referred to as inferior mesenteric vessels), ureter, reproductive vessels, Toldt's fascia, Gerota's fascia, presacral fascia, Denonvilliers' fascia, levator ani muscle, etc.
[0041] In the coordinate system of the three-dimensional model, the spatial travel path of the virtual endoscope is constructed, such as constructing a virtual endoscope curve that describes the typical laparoscopic travel route of total mesorectal excision (TME). Where [0,1] is the normalized path parameter range, Represents three-dimensional Euclidean space, parameters This represents the normalized parameter for the arc length along the path. This corresponds to the three-dimensional spatial location. Its overall approach follows the clinical "umbilical approach". Top view of the pelvis IMA / IMV area Free Toldt's fascia Holy Plane medial to lateral dissection Free on the outer side of the white line on the left Posterior three-plane free pelvic cavity The stable geometric pattern of the "deep pelvic levator anal muscle funnel area" can be understood as the funnel-shaped area formed by the levator anal muscle surrounding the distal rectum. In this invention, it is uniformly referred to as the deep pelvic levator anal muscle funnel area.
[0042] Then, the spatial travel path interval [0,1] is divided into several path sub-intervals associated with each expected clinical scenario. Each paragraph Corresponding to a clinical scenario (Such as vascular area treatment, Holy Plane dissection, posterior three-plane pelvic cavity treatment, deep pelvic levator ani infundibulum area treatment, etc.), among which, This is the index of the path sub-interval. and These represent the start and end parameters of the sub-interval on the overall path parameter interval [0,1].
[0043] like Figure 2 As shown, based on the patient's preoperative 3D CT data, a simplified anatomical model including the lower thoracic cage, lumbar spine, sacrum, pelvis, and the course of the sigmoid colon and rectum can be reconstructed within a unified volume data bounding box. This bounding box corresponds to the volume coordinate range of the 3D model and is used to define the virtual laparoscopic path curve. The space for movement is designed to ensure that the trajectory of the virtual endoscope remains within the patient's anatomical structure. In the three-dimensional model, based on the typical movement pattern of the laparoscopic lens during total mesorectal resection, a continuous and smooth virtual endoscope path curve is constructed starting from the anterior abdominal wall region of the umbilicus approach and following the course of the sigmoid mesentery and rectal mesentery.
[0044] S2. Select at least one key node in each of the path sub-intervals, and predefine the expected set of anatomical structures for each key node; integrate the spatial travel path, the path sub-intervals, the key nodes, and the set of anatomical structures to form a path prior model;
[0045] Specifically, discrete nodes are selected within each sub-interval. ,in, The center position of the lens. In the direction of the line of sight, Define the desired scene label. Then, predefine the expected visible anatomy set for each selected key node. ,in, For the index of the path node, This represents the set of indices for all anatomical structure labels in the 3D model. Indicates at node The expected subset of anatomical structures is identified (e.g., the anatomical set corresponding to the IMA root node includes the IMA, IMV, ureter, reproductive vessels, and Toldt's fascia). Then, the spatial path, the path sub-intervals, the key nodes, and the set of anatomical structures are integrated to obtain a patient-specific path prior model. .
[0046] like Figure 2As shown, the virtual laparoscopic path curve runs from top to bottom through the upper abdomen, left lower abdomen, sacral promontory plane, and down to the deep pelvic levator ani infundibulum area, with multiple key nodes discretely set along the path. to , among which, nodes Located in the pelvic cavity overhead exploration position, corresponding to the surgeon's initial overhead observation of the pelvis and rectal mesentery upon entering the abdominal cavity; node Positioned near the root of the inferior mesenteric vessels, corresponding to the typical observation angle for handling the inferior mesenteric arteries and veins and dissecting them from medial to lateral; Node Located near the left Toldt's fascia, representing the viewpoint of lateral displacement along the fascial space after entering the "Holy plane"; node With nodes Located in the upper, middle, and posterior rectal regions, respectively, these are used to characterize the observation location during dissection along the presacral space and parapelvic nerve region; nodes Located in the infundibulum region of the levator ani muscle in the deep pelvis, this represents a typical observational perspective during the dissection and resection of the distal rectum in a low or ultra-low position.
[0047] exist Figure 2 In the diagram, dashed boxes mark the neighborhood regions of each key node in 3D space, black dots represent the position of the optical center of the virtual endoscope, and arrows indicate the orientation of the visual axis of the corresponding node. By pre-coding the position, orientation, and associated anatomical regions of the nodes, a continuous virtual endoscope path prior model can be formed in the 3D model from top to bottom. When implementing this method, if the scene determination result indicates that the current surgical stage corresponds to a certain node or the path sub-interval of the node, the virtual endoscope viewpoint can be switched to the position and orientation shown by that node, thereby reproducing the observation route and field of view that matches the actual surgical stage on the preoperative 3D model.
[0048] S3. During the operation, the video sequence of the virtual endoscope acquired in real time is segmented to obtain segmentation results including anatomical structures and surgical instruments; based on the segmentation results of multiple consecutive frames within a preset time window, semantic features reflecting the dynamic changes of the surgical scene are extracted, and the current state of the surgical scene is determined according to the semantic features.
[0049] Specifically, obtaining the segmentation results, including anatomical structures and surgical instruments, includes:
[0050] During the surgery, video sequences from a virtual endoscope are acquired in real time, and each frame of the video sequence is input into an instance segmentation network for processing to obtain segmentation results for each video frame, including anatomical structures and surgical instruments.
[0051] Specifically, during the surgery, video sequences from the virtual endoscope are acquired in real time. And input each frame of the video sequence into an instance segmentation network (e.g., The data is processed to obtain segmentation results for each frame of video, including anatomical structures and surgical instruments.
[0052] Specifically, the extraction of semantic features reflecting the dynamic changes of the surgical scene based on the segmentation results of multiple consecutive frames within a preset time window includes:
[0053] Based on the segmentation results, an instance set corresponding to each video frame is formed. The instance set includes category labels, contour masks, and positional distribution of various anatomical structures and surgical instruments in the image.
[0054] Within a preset time window, statistical analysis is performed on the set of instances within the time window to extract semantic features that reflect the dynamic changes of the surgical scene, and the corresponding semantic feature vectors are obtained.
[0055] Preferably, the semantic features reflecting the dynamic changes of the surgical scene specifically include: the number of frames in which various anatomical structures and surgical instruments appear within a preset time window, the average area ratio, and at least one of the relative positions, distances, and occlusion relationships between key anatomical structures.
[0056] Specifically, based on the segmentation results of each video frame, including anatomical structures and surgical instruments, an instance set corresponding to each video frame is formed. ,in, Category labels (such as IMA, IMV, ureter, rectal mesentery, Gerota fascia, electrocautery hook, ultrasonic scalpel, etc.). This serves as a spatial mask. The instance set includes category labels, contour masks, and positional distributions of various anatomical structures and surgical instruments in the image, for subsequent semantic feature statistics.
[0057] In length of time window Within this process, statistical analysis is performed on the set of instances for each video frame to calculate the category. The statistical characteristics include the number of frames that appear. Average regional proportion Simultaneously, the spatial relationships between key anatomical structures were extracted (these spatial relationships include the relative positions, distances, and occlusion relationships between key anatomical structures, such as the centroid distance between the IMA and the ureter and reproductive vessels). relative angle Occlusion ratio (etc.), and then calculate the average or trend quantity over the time window. .in, Indicates the frame length of the time window. For the current moment The end, length is The video frame index set, Indicates the number of video frames within the time window; This indicates an indicator function; its value is 1 when its internal condition is true and 0 when the condition is false, used to determine the state of a specific video frame. Does the category exist? Examples; This represents the pixel area of the corresponding mask or image. This indicates that it belongs to a category in a specific video frame u. The segmentation mask, Represents a specific video frame The entire image.
[0058] Finally, the above statistics are used to construct a semantic feature vector within the time window. This semantic feature vector characterizes the field of vision from "abdominal cavity" within a certain period of time. Vascular area Toldt plane Pelvic three planes The dynamic semantic pattern of "deep pelvis".
[0059] Specifically, determining the current surgical scenario state based on the semantic features includes:
[0060] The obtained semantic feature vector is input into a pre-trained classification model or a preset judgment rule to determine the current scene state of the surgery.
[0061] Specifically, define a finite set of states. Each state This corresponds to a surgical scenario stage (e.g.: Corresponding to pelvic top view exploration, Corresponding IMA / IMV vascular area Corresponding to Holy Plane free, Corresponding to the posterior three planes of the pelvis (Corresponding to the treatment of the deep pelvic levator ani infundibulum area, etc.), and then reconstructing the mapping. It can be implemented as a rule-based decision function or a parameterized model, for example, through likelihood. With prior transition probability The recursive estimation is performed, and the specific formula for the recursive estimation is as follows:
[0062]
[0063] in, Represents the time from initial time 1 to the current time. The semantic feature sequence of all time windows, This represents conditional probability.
[0064] Use entry / exit thresholds / With minimum sustained frame length A hysteresis mechanism is constituted only when And maintain Scene transitions are confirmed only when the frame rate is above a certain threshold, thus allowing determination of the current surgical scenario and obtaining the current scenario state. For example, it may be determined as the treatment stage of the intestinal vascular area, the freeing stage of the "Holy Plane" fascia space under Toldt's fascia, the freeing stage of the posterior three planes of the pelvis, or the treatment stage of the levator ani muscle infundibulum area of the deep pelvis.
[0065] like Figure 4 As shown in the upper dashed box, multiple frames of laparoscopic images were continuously acquired during the surgery. Each frame was segmented, and the identified target regions, such as the inferior mesenteric artery, inferior mesenteric vein, left ureter, rectal mesentery, and Toldt's fascia, were marked with thick outlines in the image. As time progresses, the morphology, location, and exposure of these structures change in different frames. By segmenting several frames within the same time window, a set of instances of each anatomical target within that window can be obtained.
[0066] like Figure 4 As shown in the lower part, this disclosure performs statistical analysis on the segmentation results within a preset time window to construct a semantic feature table. For each type of anatomical structure, the number of frames it appears in within the window, its appearance ratio (the ratio of the number of appearance frames to the number of frames in the window), its average area percentage, and its average vertical and horizontal positions in the image coordinate system are statistically analyzed. Taking the inferior mesenteric artery as an example, it appears in 29 frames in the current window, with an appearance ratio of 96.7%, an average area percentage of 3.2%, an average vertical position of 0.42, and an average horizontal position of 0.53; the inferior mesenteric vein, left ureter, rectal mesentery, and Toldt's fascia are also given corresponding statistical features. The table on the right shows the results of scene determination based on the above semantic features: the current time window length is 30 frames, and the scene state is determined to be... (For example, the processing of the inferior mesenteric vascular area and the scene of medial to lateral freeing), the scene determination confidence is 0.91. By summarizing and extracting features from the segmentation results of multiple consecutive frames, this disclosure can robustly identify the current surgical scene at the time window scale, providing reliable semantic input for subsequent virtual laparoscopic path interval selection and node perspective switching.
[0067] S4. Based on the current scene state, determine the corresponding path sub-interval from the path prior model, and update the propulsion parameters of the virtual cavity mirror on the spatial travel path accordingly; when the propulsion parameters approach a certain key node predefined in the path prior model, trigger a viewpoint switch and switch the viewpoint of the virtual cavity mirror to the pose corresponding to the key node.
[0068] Specifically, updating the propulsion parameters of the virtual cavity mirror along the spatial travel path includes:
[0069] Based on the current scene state obtained from the determination, the corresponding path sub-interval is determined from the path prior model, and the path segment that matches the current scene is selected from the path sub-interval;
[0070] The path parameters of the virtual cavity mirror are updated according to the path segment, so that the virtual cavity mirror advances along the path segment.
[0071] Specifically, in order to advance the digital twin path, a mapping from scenario to path interval is defined. And typical representative mappings from scene to node. Given the current decision scenario. First, select the corresponding path sub-interval. And based on the strength of the evidence Alternatively, update the path parameters on the surgical timeline:
[0072] in This indicates that the current moment belongs to the scene. confidence level Step size factor To limit The propagation function; when When, trigger node The perspective switch sets the position of the virtual cavity mirror. The view axis is set to Parameters such as field of view are based on and Adjustments were made to present a virtual field of view that matches the current scene, among which, Represents a node In the path The corresponding normalized parameter value, The preset trigger threshold is set when the current path parameters of the virtual cavity mirror... With nodes A node switch is triggered when the parameters are sufficiently close. For a scenario from... Towards The conversion, through Interval jumps and nodes The switching enables a phased digital twin transition from the abdominal cavity to the vascular region, from the vascular region to the Toldt plane, and then to the pelvic cavity and deep pelvis. If the entropy value... If the set threshold is exceeded, the system enters "freeze / rollback" mode. Lock or backtrack to the nearest high-confidence node to avoid erroneous advancements under low-confidence conditions and retain the human perspective overlay interface.
[0073] like Figure 3 As shown, Figure 3 The left side shows a magnified section of the virtual laparoscopic path in the preoperative 3D model. The black curve represents the virtual laparoscopic path running along the sigmoid mesentery and rectal mesentery, and the black dots on this path represent selected path nodes. At the node, a camera icon and a field of view cone are overlaid to indicate the spatial position and viewing direction of the virtual cavity mirror at that node. Therefore, the node... It not only falls on a pre-constructed path curve, but also has a defined visual axis orientation to simulate the typical observation angle of the intraoperative camera in the inferior mesenteric vascular region.
[0074] Figure 3 The right side provides a table example of nodes. The semantic encoding content is as follows: The table first provides the node identifier (id), which uniquely identifies the node within the path node set; the corresponding path segment. Mark the node as located in " "Inferior mesenteric vessel segment" indicates that it is located within the inferior mesenteric vessel processing and surrounding free path area; path parameters For this node along the entire path curve The normalized arc length position, for example, taking a value of 0.23. Three-dimensional position coordinates. Here are the coordinates of the optical center of the node in the 3D model coordinate system, for example, ((12.4, -35.7, 58.2)); and the view axis direction vector. This indicates the orientation of the virtual cavity mirror, for example, ((0.35, -0.78, -0.52)), in conjunction with the field of view. (e.g., 75°) and camera roll angle (e.g., 0°) together define the imaging field of view pose of this node.
[0075] At the semantic level, the table uses corresponding scene tags. Specify node The representative surgical scenario is "treatment of the root of the inferior mesenteric vessel and freeing the entrance from the medial to the lateral side," and is achieved through the assembly of anticipated anatomical structures. The main anatomical structures that should appear in the field of view of this node are given, such as the inferior mesenteric artery, inferior mesenteric vein, left ureter, Toldt's fascia, and root of the rectal mesentery. Through the joint encoding of the above geometric parameters and semantic attributes, each path node becomes a "location". Orientation Scene In implementing this method, the scene state output by the scene determination unit 33 can be matched with the label and anatomical set of the corresponding node, thereby driving the virtual laparoscopic viewpoint to automatically switch to the node position and perspective that is most consistent with the current real surgical field of view.
[0076] Furthermore, the virtual endoscopic digital twin method can be encapsulated in a surgical navigation system. For example, the front end can use C++ / Qt and a graphics rendering library to achieve 3D visualization, and the back end can use a GPU-accelerated framework to implement neural network inference to meet real-time requirements (e.g., the frame rate can reach approximately 60 frames per second).
[0077] On the other hand, such as Figure 5 and Figure 6 As shown, this embodiment of the invention also provides a virtual endoscope digital twin device, employing the aforementioned virtual endoscope digital twin method. The device includes a data input interface 1, an image acquisition unit 2, and a processor 3. The processor 3 includes a 3D reconstruction and path encoding unit 31, an instance segmentation and semantic feature extraction unit 32, a scene determination unit 33, and a path advancement and view control unit 34, wherein:
[0078] Data input interface 1 is used to acquire the patient's preoperative medical imaging data and send the imaging data to the three-dimensional reconstruction and path coding module;
[0079] Image acquisition device 2 is used to acquire video sequences of the virtual cavity mirror in real time and send the video sequences to the instance segmentation and semantic feature extraction module;
[0080] The three-dimensional reconstruction and path coding unit 31 is used to construct a three-dimensional anatomical model of the patient based on the received image data; it is also used to construct a spatial travel path of the virtual endoscope, divide the path into sub-intervals and set several key nodes in the coordinate system of the three-dimensional anatomical model to form a path prior model.
[0081] The instance segmentation and semantic feature extraction unit 32 is used to perform instance segmentation on the received video sequence to obtain segmentation results including anatomical structures and surgical instruments; it is also used to extract semantic features reflecting the dynamic changes of the surgical scene based on the segmentation results of multiple consecutive frames within a preset time window.
[0082] The scene determination unit 33 is used to determine the current scene state of the surgery based on the semantic features sent by the instance segmentation and semantic feature extraction module.
[0083] The path advancement and view control unit 34 is used to determine the corresponding path sub-interval from the path prior model based on the current scene state sent by the scene determination module, and update the advancement parameters of the virtual cavity mirror on the spatial travel path accordingly; it is also used to trigger view switching when the advancement parameters approach a certain key node predefined in the path prior model, and switch the view of the virtual cavity mirror to the pose corresponding to the key node.
[0084] Furthermore, the device also includes a memory 4 and a display terminal 5, wherein:
[0085] The memory 4 is used to store the patient's preoperative medical imaging data, three-dimensional anatomical model, path prior model, semantic features, current scene status, key nodes and viewpoint switching sequence;
[0086] The display terminal 5 is used to simultaneously display the real intraoperative laparoscopic view and the virtual laparoscopic view generated based on the patient's three-dimensional anatomical model and path prior model.
[0087] In this embodiment, the device can be integrated into a computer terminal or dedicated workstation, including a data input interface 1, an image acquisition unit 2, a processor 3, a memory 4, and a display terminal 5. These components are interconnected via a system bus 6 or other wired / wireless communication methods. The image processor 7 is electrically connected to the image acquisition unit 2 as an external device, used to output intraoperative laparoscopic image signals to the image acquisition unit 2, and transmit the acquired real-time video data to the processor 3 through the image acquisition unit 2. The processor 3 may include one or more central processing units (CPU), graphics processing units (GPU), digital signal processors (DSP), or programmable logic devices, etc., used to execute program instructions stored in the memory 4. The memory 4 may include read-only memory 4, random access memory 4, and non-volatile storage media, etc., used to store the operating system, application programs, 3D model data, path prior data, and intraoperative feature and log information.
[0088] The processor 3 includes a 3D reconstruction and path encoding unit 31, an instance segmentation and semantic feature extraction unit 32, a scene determination unit 33, a path advancement and view control unit 34, and a human-computer interaction and log unit 35. The 3D reconstruction and path encoding unit 31 performs multi-structure segmentation and 3D reconstruction based on the patient's preoperative medical imaging data (such as CT data) to obtain a 3D model of the patient. Furthermore, a virtual endoscope travel path is constructed in the model coordinate system, the path is divided into sub-intervals, and multiple key nodes are set, thereby forming a path prior model. 3D model and path prior model This can be provided to the subsequent path advancement and perspective control unit 34 and the display terminal 5 to generate a virtual laparoscopic view; the instance segmentation and semantic feature extraction unit 32 performs instance segmentation on the intraoperative real-time laparoscopic video stream, identifies multiple target objects including blood vessels, fascia, organ surfaces, and surgical instruments, and statistically analyzes the number of frames, area proportion, and relative spatial relationship of various targets within a preset time window to generate semantic features. And output to the scene determination unit 33; the scene determination unit 33 is based on semantic features The current surgical scenario is determined according to preset rules or probability models to obtain a determination scenario representing the surgical stage or operating environment. and the determination scenario The data is sent to the path advancement and view control unit 34; the path advancement and view control unit 34 is based on the path prior model. Under the constraints of the judgment scenario Select the corresponding path sub-intervals and key nodes, update the path parameters, and generate node switching signals. and perspective switching signal The virtual laparoscopic viewpoint is advanced along a preset path to a node position and orientation that matches the current scene, thereby driving the synchronous update of the virtual 3D view; the display terminal 5 simultaneously displays the real intraoperative laparoscopic image and the patient-based 3D model. and path prior model The generated virtual laparoscopic view enables digital twin visualization of the surgical procedure; the memory 4 stores the patient's medical image data and 3D model. Path prior model semantic features Determine the scenario The system also includes information such as node and perspective switching sequences to facilitate postoperative playback, model updates, and algorithm optimization. The human-computer interaction and logging unit 35 drives the display terminal 5 to display its interface, receive user input, record scene judgment and perspective switching history, and upload relevant logs through the network interface module 8.
[0089] Furthermore, those skilled in the art will understand that the 3D reconstruction and path encoding unit 31, the instance segmentation and semantic feature extraction unit 32, the scene determination unit 33, the path advancement and view control unit 34, and the human-computer interaction and log unit 35 can all be stored in the memory 4 as software modules and implemented by the processor 3 during execution, or they can be implemented through hardware logic or a combination of hardware and software, which does not constitute a limitation of this disclosure.
[0090] In another aspect, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by processor 3, implements the aforementioned virtual cavity mirror digital twin method.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for creating a virtual laparoscopic digital twin, characterized in that, The method includes the following steps: Based on the acquired preoperative medical imaging data of the patient, a three-dimensional anatomical model of the patient is constructed; in the coordinate system of the three-dimensional anatomical model, a spatial travel path of the virtual endoscope is constructed, and the spatial travel path is divided into several path sub-intervals associated with each expected clinical scenario; At least one key node is selected within each of the path sub-intervals, and a set of anatomical structures expected to be visible for each key node is predefined; the spatial travel path, the path sub-intervals, the key nodes, and the set of anatomical structures are integrated to form a path prior model; During the surgery, the video sequence of the virtual endoscope acquired in real time is segmented to obtain segmentation results including anatomical structures and surgical instruments; based on the segmentation results of multiple consecutive frames within a preset time window, semantic features reflecting the dynamic changes of the surgical scene are extracted, and the current state of the surgical scene is determined according to the semantic features. Based on the current scene state, the corresponding path sub-interval is determined from the path prior model, and the propulsion parameters of the virtual cavity mirror on the spatial travel path are updated accordingly. When the propulsion parameters approach a certain key node predefined in the path prior model, a viewpoint switch is triggered, and the viewpoint of the virtual cavity mirror is switched to the pose corresponding to the key node.
2. The virtual endoscope digital twin method as described in claim 1, characterized in that, The construction of a three-dimensional anatomical model of the patient based on preoperative medical imaging data specifically includes: The patient's preoperative medical imaging data is acquired, and the medical imaging data is segmented into multiple anatomical structures to obtain a three-dimensional model including several anatomical labels. The medical imaging data includes three-dimensional CT data.
3. The virtual endoscope digital twin method as described in claim 1, characterized in that, The obtained segmentation results, including anatomical structures and surgical instruments, specifically include: During the surgery, video sequences from a virtual endoscope are acquired in real time, and each frame of the video sequence is input into an instance segmentation network for processing to obtain segmentation results for each video frame, including anatomical structures and surgical instruments.
4. The virtual endoscopic digital twin method as described in claim 3, characterized in that, The segmentation results based on multiple consecutive frames within a preset time window are used to extract semantic features reflecting the dynamic changes of the surgical scene, specifically including: Based on the segmentation results, an instance set corresponding to each video frame is formed. The instance set includes category labels, contour masks, and positional distribution of various anatomical structures and surgical instruments in the image. Within a preset time window, statistical analysis is performed on the set of instances within the time window to extract semantic features that reflect the dynamic changes of the surgical scene, and the corresponding semantic feature vectors are obtained.
5. The virtual endoscopic digital twin method as described in claim 4, characterized in that, The semantic features reflecting the dynamic changes of the surgical scene specifically include: the number of frames in which various anatomical structures and surgical instruments appear within a preset time window, the average area ratio, and at least one of the relative positions, distances, and occlusion relationships between key anatomical structures.
6. The virtual endoscopic digital twin method as described in claim 4, characterized in that, The step of determining the current surgical scenario state based on the semantic features specifically includes: The obtained semantic feature vector is input into a pre-trained classification model or a preset judgment rule to determine the current scene state of the surgery.
7. The virtual endoscopic digital twin method as described in claim 1, characterized in that, The updating of the propulsion parameters of the virtual cavity mirror along the spatial travel path specifically includes: Based on the current scene state obtained from the determination, the corresponding path sub-interval is determined from the path prior model, and the path segment that matches the current scene is selected from the path sub-interval; The path parameters of the virtual cavity mirror are updated according to the path segment, so that the virtual cavity mirror advances along the path segment.
8. A virtual laparoscopic digital twin device, employing the virtual laparoscopic digital twin method according to any one of claims 1-7, characterized in that, The device includes a data input interface, an image acquisition unit, and a processor. The processor includes a 3D reconstruction and path encoding unit, an instance segmentation and semantic feature extraction unit, a scene determination unit, and a path advancement and view control unit, wherein: The data input interface is used to acquire the patient's preoperative medical imaging data and send the imaging data to the three-dimensional reconstruction and path coding module; The image acquisition device is used to acquire video sequences of the virtual cavity mirror in real time and send the video sequences to the instance segmentation and semantic feature extraction module; The three-dimensional reconstruction and path coding unit is used to construct a three-dimensional anatomical model of the patient based on the received image data; it is also used to construct the spatial travel path of the virtual endoscope, divide the path into sub-intervals and set several key nodes in the coordinate system of the three-dimensional anatomical model to form a path prior model. The instance segmentation and semantic feature extraction unit is used to perform instance segmentation on the received video sequence to obtain segmentation results including anatomical structures and surgical instruments; it is also used to extract semantic features reflecting the dynamic changes of the surgical scene based on the segmentation results of multiple consecutive frames within a preset time window. The scene determination unit is used to determine the current scene state of the surgery based on the semantic features sent by the instance segmentation and semantic feature extraction module. The path advancement and view control unit is used to determine the corresponding path sub-interval from the path prior model based on the current scene state sent by the scene determination module, and update the advancement parameters of the virtual cavity mirror on the spatial travel path accordingly; it is also used to trigger view switching when the advancement parameters approach a certain key node predefined in the path prior model, and switch the view of the virtual cavity mirror to the pose corresponding to the key node.
9. A virtual laparoscopic digital twin device as described in claim 8, characterized in that, The device further includes a memory and a display terminal, wherein: The memory is used to store the patient's preoperative medical imaging data, three-dimensional anatomical model, path prior model, semantic features, current scene state, key nodes and viewpoint switching sequence; The display terminal is used to simultaneously display the actual intraoperative laparoscopic view and a virtual laparoscopic view generated based on the patient's three-dimensional anatomical model and path prior model.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements a virtual endoscopic digital twin method according to any one of claims 1-7.