Urinary calculus robot assisting method and device, electronic equipment and storage medium
By performing temporal alignment and spatial registration of multimodal data on the urinary stone robot, the optimal puncture path and motion vector are generated, solving the problem of information fragmentation and improving the safety and predictability of urinary stone surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-03-13
AI Technical Summary
Existing urinary stone robots have the risks of information fragmentation and 'blind insertion' during surgery. They rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. They also lack the ability to effectively integrate and register preoperative multimodal data with intraoperative real-time video stream data.
The Kalman filter algorithm is used to perform time alignment of multimodal data, and the image registration energy function and deformation field model are combined to perform spatial registration, generating the optimal puncture path and the corresponding optimal motion vector to assist the urinary stone robot in performing surgical operations.
It achieves precise integration of preoperative and intraoperative information, reduces reliance on doctors' experience, improves the safety and predictability of urinary stone surgery, and reduces the risk of 'blind puncture'.
Smart Images

Figure CN121647818A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of surgical robot-assisted decision-making, and more specifically, to a robot-assisted method, device, electronic device, and storage medium for urinary calculi treatment. Background Technology
[0002] Urinary tract stones are a common disease in urology, and their diagnosis and treatment typically include preoperative imaging diagnosis, surgical treatment, and postoperative stone composition analysis. Current mainstream treatment methods suffer from information fragmentation and the risk of "blind puncture." Preoperative CT images and intraoperative real-time endoscopic videos exist in different spatiotemporal coordinate systems and cannot be precisely aligned. When planning the puncture path, the surgeon needs to mentally map the "two-dimensional image to three-dimensional space," making puncture path planning heavily reliant on the surgeon's personal experience. Due to the lack of precise real-time guidance, the surgeon struggles to accurately determine the position and direction of the puncture needle during the procedure, increasing the risk of damage to surrounding tissues such as the renal pelvis, blood vessels, and intestines. This information fragmentation not only reduces surgical precision but also prolongs the operation time and may lead to postoperative complications.
[0003] Furthermore, existing urological stone surgery robots are mostly "blind arms," lacking intelligent decision-making capabilities based on real-time multimodal information. Surgical outcomes heavily rely on the surgeon's skill and experience, making standardization and widespread adoption difficult. During urological stone surgery, the lack of technology to effectively integrate and register preoperative multimodal data with intraoperative real-time video stream data prevents the robot from obtaining accurate real-time guidance information, thus limiting its application in complex surgeries. This lack of information integration and registration capabilities prevents the robot from making intelligent decisions based on the patient's specific condition and real-time feedback, further exacerbating the risk of "blind puncture" and hindering the intelligent development of urological stone surgery robots.
[0004] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0005] The purpose of this application is to provide a robot-assisted method, device, electronic device, and storage medium for urinary stone surgery. By effectively integrating and registering preoperative multimodal data with real-time intraoperative video stream data, the data is input into a preset path generation model and a preset motion generation model to generate the optimal puncture path and the corresponding optimal motion vector. This assists the urinary stone robot in performing surgical operations, solving the problems of information fragmentation and "blind puncture" risks in existing urinary stone robots during surgery, which rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. This application can accurately integrate preoperative and intraoperative information to reduce reliance on the doctor's experience and improve the safety and predictability of urinary stone surgery.
[0006] In a first aspect, this application provides a robot-assisted method for treating urinary stones, comprising: S1, acquire multimodal data of the patient before surgery; S2, The Kalman filter algorithm is used to perform time alignment on the multimodal data to obtain time-aligned multimodal data; S3, using a preset image registration energy function and a preset deformation field model, spatially register the real-time video stream data of the patient during surgery with the time-aligned multimodal data to obtain the registered real-time modal data; S4, input the registered real-time modal data into the preset path generation model and the preset action generation model to generate the optimal puncture path and the corresponding optimal action vector; S5, output the optimal puncture path and the optimal motion vector to assist the urinary stone robot in performing surgical operations.
[0007] The urological stone robot-assisted method provided in this application can assist the urological stone robot by effectively integrating and registering preoperative multimodal data with intraoperative real-time video stream data, and then inputting them into a preset path generation model and a preset motion generation model to generate the optimal puncture path and the corresponding optimal motion vector. This assists the urological stone robot in performing surgical operations, solving the problems of information fragmentation and "blind puncture" risks in existing urological stone robots during surgery, which rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. It can accurately integrate preoperative and intraoperative information to reduce reliance on the doctor's experience and improve the safety and predictability of urological stone surgery.
[0008] Optionally, the real-time video stream data of the patient during surgery and the time-aligned multimodal data are spatially registered using a preset image registration energy function and a preset deformation field model to obtain registered real-time modal data, including: Real-time video stream data of the patient during surgery is acquired. The real-time video stream data is compensated using a preset deformation field model to obtain compensated real-time video stream data. Using a preset image registration energy function, spatial registration is performed on the compensated real-time video stream data and the time-aligned multimodal data to obtain registered real-time modal data.
[0009] The urinary calculi robot-assisted method provided in this application can assist the urinary calculi robot. By introducing a deformation field model to compensate for real-time video stream data, and introducing an image registration energy function to spatially register the compensated real-time video stream data and time-aligned multimodal data, it can effectively handle tissue deformation that may occur during surgery, further improve the accuracy and robustness of image registration, and make real-time guidance information more reliable.
[0010] Optionally, the construction steps of the preset image registration energy function include: Based on the similarity of the images, an initial image registration energy function is constructed; By introducing a deformation field regularization term, the initial image registration energy function is optimized to obtain the preset image registration energy function.
[0011] Optionally, the registered real-time modal data is input into a preset path generation model and a preset motion generation model to generate an optimal puncture path and a corresponding optimal motion vector, including: The registered real-time modal data is input into a preset feature enhancement model to obtain a stone feature enhancement image; The stone composition probability vector and stone hardness estimation vector are extracted from the registered real-time modal data. The enhanced image of the stone features, the probability vector of the stone composition, and the estimation vector of the stone hardness are input into a preset path generation model and a preset action generation model to generate the optimal puncture path and the corresponding optimal action vector.
[0012] The urinary stone robot-assisted method provided in this application can assist a urinary stone robot. By extracting the stone composition probability vector and stone hardness estimation vector from multimodal data, the method not only considers real-time modal data when generating the optimal puncture path and action vector, but also combines the stone composition and hardness information, making the generated path and action more in line with the actual surgical needs and improving the personalization and effectiveness of the surgery.
[0013] Optionally, before inputting the registered real-time modal data into a preset feature enhancement model to obtain the stone feature enhancement image, the method further includes: Using a pre-trained encoder, the registered real-time modal data and the pose information of the urinary stone robot are mapped into a common space, so that multiple data at the same surgical moment are located in the same common space.
[0014] Optionally, the enhanced image of the stone features, the probability vector of the stone composition, and the estimation vector of the stone hardness are input into a preset path generation model and a preset action generation model to generate an optimal puncture path and a corresponding optimal action vector, including: The enhanced image of the stone features is input into a preset path generation model to generate the optimal puncture path; The optimal puncture path, the enhanced image of the stone features, the probability vector of the stone composition, and the estimation vector of the stone hardness are input into a preset action generation model to generate the optimal action vector corresponding to the optimal puncture path.
[0015] Optionally, after outputting the optimal puncture path and the optimal motion vector to assist the urinary stone robot in performing surgical operations, the process includes: Obtain surgical outcome data after a urinary calculi procedure performed by a robotic surgical device; The multimodal data and the surgical result data are compared to determine whether the surgery is completed. If yes, the surgery is confirmed to be completed. If no, the modal data corresponding to the surgical result data is determined as the new registered real-time modal data, and the process returns to step S3.
[0016] Secondly, this application provides a robotic-assisted device for urinary stones, comprising: The acquisition module is used to acquire multimodal data of patients before surgery; The alignment module is used to perform time alignment on the multimodal data using a Kalman filter algorithm to obtain time-aligned multimodal data. The registration module is used to spatially register the real-time video stream data of the patient during surgery with the time-aligned multimodal data using a preset image registration energy function and a preset deformation field model, so as to obtain the registered real-time modal data. The generation module is used to input the registered real-time modal data into a preset path generation model and a preset action generation model to generate the optimal puncture path and the corresponding optimal action vector. The output module is used to output the optimal puncture path and the optimal motion vector to assist the urinary stone robot in performing surgical operations.
[0017] This robotic-assisted device for urinary stone surgery effectively integrates and registers preoperative multimodal data with real-time intraoperative video stream data, then inputs it into a preset path generation model and a preset motion generation model to generate the optimal puncture path and corresponding optimal motion vector. This assists the urinary stone robot in performing surgical operations, solving the problems of information fragmentation and "blind puncture" risks in existing urinary stone robots during surgery, which rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. It can accurately fuse preoperative and intraoperative information to reduce reliance on the doctor's experience and improve the safety and predictability of urinary stone surgery.
[0018] Thirdly, this application provides an electronic device including a processor and a memory, the memory storing a computer program executable by the processor, wherein when the processor executes the computer program, it performs the steps of the robot-assisted method for urinary stones as described above.
[0019] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the robot-assisted method for urinary stones described above.
[0020] Beneficial effects: The urinary stone robot-assisted method, device, electronic device, and storage medium provided in this application effectively integrate and register preoperative multimodal data with intraoperative real-time video stream data, and input them into a preset path generation model and a preset motion generation model to generate the optimal puncture path and the corresponding optimal motion vector. This assists the urinary stone robot in performing surgical operations, solving the problems of information fragmentation and "blind puncture" risks in existing urinary stone robots during surgery, which rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. It can accurately integrate preoperative and intraoperative information to reduce reliance on the doctor's experience and improve the safety and predictability of urinary stone surgery. Attached Figure Description
[0021] Figure 1 A flowchart of a robot-assisted method for urinary stone treatment provided in an embodiment of this application.
[0022] Figure 2 This is a schematic diagram of the structure of the urinary stone robot-assisted device provided in the embodiments of this application.
[0023] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0024] Labeling Explanation: 1. Acquisition Module; 2. Alignment Module; 3. Registration Module; 4. Generation Module; 5. Output Module; 301. Processor; 302. Memory; 303. Communication Bus. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0026] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0027] Please refer to Figure 1 , Figure 1 This application discloses a robot-assisted method for treating urinary stones, which is used to assist a urinary stone treatment robot and includes the following steps: Step S1: Obtain the patient's multimodal data before surgery; Step S2: Use the Kalman filter algorithm to perform time alignment on the multimodal data to obtain time-aligned multimodal data; Step S3: Using a preset image registration energy function and a preset deformation field model, spatially register the real-time video stream data of the patient during surgery with the time-aligned multimodal data to obtain the registered real-time modal data. Step S4: Input the registered real-time modal data into the preset path generation model and the preset motion generation model to generate the optimal puncture path and the corresponding optimal motion vector. Step S5: Output the optimal puncture path and optimal motion vector to assist the urinary stone robot in performing surgical operations.
[0028] This robot-assisted method for urinary stone surgery effectively integrates and registers preoperative multimodal data with real-time intraoperative video stream data, then inputs it into a preset path generation model and a preset motion generation model to generate the optimal puncture path and corresponding optimal motion vector. This assists the urinary stone robot in performing surgical operations, solving the problems of information fragmentation and "blind puncture" risks in existing urinary stone robots during surgery, which rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. It can accurately integrate preoperative and intraoperative information to reduce reliance on the doctor's experience and improve the safety and predictability of urinary stone surgery.
[0029] Specifically, in step S1, multimodal data of the patient before surgery is acquired using existing technology. The multimodal data includes preoperative CT scan images, MRI images, lithotripsy composition analysis results (calculated based on Fourier transform infrared spectroscopy (FTIR)) and corresponding stone composition probability vectors (the stone composition probability vector represents the probability distribution of different stone components, which can be obtained by inputting the lithotripsy composition analysis results into a pre-set composition probability calculation model to probabilistically evaluate the chemical composition of the stone (such as calcium oxalate, uric acid, magnesium ammonium phosphate, etc.). This composition probability calculation model is a trained machine learning algorithm model) and stone hardness estimation vectors (the stone hardness estimation vector represents the quantitative value or range of stone hardness, which can be obtained by inputting the lithotripsy composition analysis results into a pre-set stone hardness estimation model to estimate the physical hardness of the stone. This stone hardness estimation model is a trained machine learning algorithm model).
[0030] Specifically, in step S2, a hardware timestamp is generated using a precision clock, and the offset between the master and slave clocks is optimally estimated using a Kalman filter to provide a unified time reference with an error of less than 1ms for multimodal data. This allows for time alignment of the multimodal data, resulting in time-aligned multimodal data. This eliminates data inconsistencies caused by differences in acquisition time and ensures that all preoperative data maintain consistency in the time dimension.
[0031] The state equation for the Kalman filter is as follows: ; in, Let k be the state vector (i.e., clock offset). Let k be the state vector at time k-1; Here is the state transition matrix. , The sampling period; This is process noise, reflecting the rate of change of the clock offset, and can be set according to actual needs.
[0032] The observation equation for Kalman filtering is as follows: ; in, Let k be the observation vector (i.e., the actual time data). For the observation matrix, ; To observe noise, settings can be adjusted based on the calculation error of the time data.
[0033] The state equation and observation equation are iterated using Kalman filtering: ; ; ; in, The Kalman gain at time k; Let be the error covariance matrix at time k-1; Let be the error covariance matrix at time k; The predicted clock offset at time k; This is the predicted clock offset at time k-1.
[0034] In summary, a unified time base is calculated to align the time of multimodal data, resulting in time-aligned multimodal data.
[0035] Specifically, in step S3, the real-time video stream data of the patient during surgery and the time-aligned multimodal data are spatially registered using a preset image registration energy function and a preset deformation field model to obtain registered real-time modal data, including: Real-time acquisition of video stream data of patients during surgery; The real-time video stream data is compensated by a preset deformation field model to obtain compensated real-time video stream data. Using a preset image registration energy function, spatial registration is performed on the compensated real-time video stream data and the time-aligned multimodal data to obtain the registered real-time modal data.
[0036] In step S3, during the robot-assisted procedure for urinary stone treatment, video image sequences of the patient's surgical area are continuously acquired using an endoscope or other real-time imaging equipment to form a continuous sequence of image frames, resulting in real-time video stream data. This video stream data reflects the dynamic changes in the surgical field and forms the basis for real-time control and operation.
[0037] Due to respiratory movements during surgery, the urinary system experiences rhythmic displacement, which can cause dynamic deformation and interfere with subsequent registration. Therefore, a deformation field model is designed to address this issue, ensuring that the compensated real-time video stream data more accurately reflects the true anatomical structure of the surgical area, thus reducing the interference of dynamic deformation on subsequent registration. The pre-designed deformation field model is as follows: ; in, For the compensated surgical time (i.e., in the compensated real-time video stream data) The spatial coordinates of point m after the deformation field has been applied; For the point during surgery (i.e., in the real-time video stream data) The spatial coordinates of point m, that is, the spatial coordinates of point m after it has been subjected to the deformation field. This indicates that static deformation due to respiratory movements is not considered. , For points before surgery (i.e., in time-aligned multimodal data) spatial coordinates, For point m (to point The displacement of ) The spatial distribution of the effects of respiratory motion can be learned from 4D-CT (four-dimensional CT scan images) through principal component analysis (PCA); It is a time function of the respiratory signal and can be obtained by real-time observation of thoracic impedance, vPPG (virtual portal vein pressure gradient), or external markers.
[0038] Using a preset image registration energy function, spatial registration is performed on the compensated real-time video stream data and the time-aligned multimodal data. The optimal registration transformation is found by optimizing the image registration energy function to obtain the registered real-time modal data.
[0039] Specifically, the steps for constructing the preset image registration energy function include: Based on the similarity of the images, an initial image registration energy function is constructed; By introducing a deformation field regularization term, the initial image registration energy function is optimized to obtain a preset image registration energy function.
[0040] When constructing the preset image registration energy function, an initial image registration energy function is obtained based on the image similarity (such as the similarity between corresponding CT scan images) in real-time video stream data and multimodal data. Specifically, the initial image registration energy function is as follows: ; in, For optimal deformation field; Register the energy function for the initial image; Let this be the initial objective function; This is a similarity metric used to calculate similarity. Preoperative 3D grayscale CT images in DICOM format obtained by a CT scanner are used to provide high-resolution anatomical structures. for Image midpoint Position coordinates; This method allows for the real-time acquisition of two-dimensional / three-dimensional grayscale images of anatomical structures via ultrasound probes or endoscopic cameras during surgery. Intraoperative ultrasound / endoscopic images The coordinates of the midpoint m; , It is a deep feature extractor, which extracts high-level semantic features from images. The square of the L2 norm, The difference between two feature vectors is measured by calculating the Euclidean distance between the corresponding feature vectors of the CT scan image and the ultrasound / endoscopic image; the initial image registration energy function. This indicates that the initial objective function is... When the minimum value is obtained, and The position makes the optimal deformation field To achieve the optimal state, i.e., when the image is most similar, determine the optimal deformation field at this point. It can achieve registration.
[0041] The purpose of introducing a deformation field regularization term is to constrain the smoothness and physical rationality of the deformation field, avoiding irregular or discontinuous jumps, thereby improving the robustness and accuracy of the registration results. Specifically, the deformation field regularization term is constrained using a viscous fluid model, which is defined as follows: ; in, For deformation field regularization term; It is a Laplace operator, which can calculate the curvature of the deformation field using the finite difference method; Let i be the i-th displacement component. , This represents the first component of the displacement, that is, the displacement in the horizontal direction. This is the second displacement component, that is, the displacement in the vertical direction. The third displacement component is the displacement in the vertical direction, where x is the abscissa of point m, y is the ordinate of point m, and z is the ordinate of point m.
[0042] In summary, the preset image registration energy function is as follows: .
[0043] Therefore, a deformation field regularization term is introduced into the initial image registration energy function to obtain a preset image registration energy function. This imposes constraints on the smoothness and continuity of the deformation field, ensuring that during the optimization process, not only image similarity but also the physical rationality of the deformation field must be considered. This combined approach allows the registration process to maintain image matching accuracy while effectively suppressing the ill-conditioned behavior of the deformation field, thereby obtaining more stable and reliable registration results.
[0044] Specifically, in step S4, the registered real-time modal data is input into a preset path generation model and a preset motion generation model to generate the optimal puncture path and the corresponding optimal motion vector, including: The registered real-time modal data is input into the preset feature enhancement model to obtain the stone feature enhancement image; The probability vector of stone composition and the estimation vector of stone hardness are extracted from the registered real-time modal data. The enhanced image of the stone features, the probability vector of the stone composition, and the estimation vector of the stone hardness are input into the preset path generation model and the preset action generation model to generate the optimal puncture path and the corresponding optimal action vector.
[0045] In step S4, the registered real-time modal data is input into a preset feature enhancement model to highlight the features of the stone region, suppress noise, and enhance the contrast or texture information of the image, resulting in a stone feature-enhanced image. This provides clearer and more discriminative visual information for subsequent path and action generation. Feature enhancement makes key information such as the stone's boundaries and internal structure more apparent, helping the preset path generation model and action generation model to more accurately identify and locate the stone. The feature enhancement model is a deep learning model, such as a convolutional neural network (CNN) or a generative adversarial network (GAN). Deep learning models are existing technologies and will not be detailed here.
[0046] From the registered real-time modal data, the probability vector of stone composition and the estimation vector of stone hardness are extracted. These vectors provide information on the intrinsic properties of the stone, which is crucial for selecting appropriate treatment parameters such as laser power and frequency of the urinary stone robot.
[0047] Specifically, in step S4, before inputting the registered real-time modal data into the preset feature enhancement model to obtain the stone feature enhancement image, the following steps are also included: Using a pre-trained encoder, the registered real-time modal data and the pose information of the urinary stone robot are mapped into a common space, so that multiple data at the same surgical moment are located in the same common space.
[0048] In step S4, before inputting the registered real-time modal data into the preset feature enhancement model, a pre-trained encoder can be used to map the registered real-time modal data and the pose information of the urinary stone robot into a 512-dimensional common space. For example, ResNet-50 (or other CNNs) can be used to extract image features from the registered real-time modal data, and then projected to 512 dimensions through a fully connected layer; 1D-CNN can be used to extract features from the lithotripsy component analysis results in the registered real-time modal data, and then projected to 512 dimensions; MLP (Multilayer Perceptron) can be used to extract the pose features of the urinary stone robot and project them to 512 dimensions. By mapping high-dimensional, heterogeneous real-time modal data and the pose information of the urinary stone robot into the common space through a pre-trained encoder, the aim is to reduce the representation distance between different modalities of the same sample, making multiple data at the same surgical moment closer in the common space, thereby effectively reducing the dimensionality of the data, removing redundant information, and bridging the semantic gap between different data sources. In the public space, these data are uniformly represented as low-dimensional, semantically consistent feature vectors, enabling subsequent path generation and action generation models to extract key features more efficiently and accurately.
[0049] Specifically, in step S4, the feature-enhanced image, the stone composition probability vector, and the stone hardness estimation vector are input into a preset path generation model and a preset action generation model to generate the optimal puncture path and the corresponding optimal action vector, including: The feature-enhanced image is input into a preset path generation model to generate the optimal puncture path; The optimal puncture path, feature-enhanced image, stone composition probability vector, and stone hardness estimation vector are input into the preset action generation model to generate the optimal action vector corresponding to the optimal puncture path.
[0050] In step S4, the enhanced image is input into a preset path generation model. Based on the spatial information and pathological features provided by the enhanced image, the model plans an optimal puncture path from the urinary stone robot's operating point to the urinary stone, avoiding critical tissues, blood vessels, or nerves, and ensuring the safety and effectiveness of the puncture procedure. After generating the optimal puncture path, it must be reviewed by a doctor before proceeding to the next step. The path generation model is a model built by combining existing path search algorithms (such as Dijkstra's algorithm, A* algorithm, or RRT fast random tree algorithm) and existing deep learning algorithms (such as neural network algorithms). It can generate a safe, efficient, and executable virtual puncture channel from the skin entry point to the target stone based on both static and dynamic anatomical environments.
[0051] The optimal puncture path, along with the enhanced image, stone composition probability vector, and stone hardness estimation vector, are input into a pre-defined motion generation model. Based on this comprehensive information, the model generates an optimal motion vector that highly matches the optimal puncture path. This optimal motion vector represents a series of precise instructions required for the urinary stone robot to perform the puncture operation, including optimal puncture depth, optimal puncture angle, optimal laser power, optimal frequency, and optimal perfusion pressure. The motion generation model, constructed using existing reinforcement learning algorithms (such as Q-learning), transforms the planned path (optimal puncture path) into specific, precise motion commands and energy parameters for the end effector (puncture needle / laser fiber) of the urinary stone robot at each moment.
[0052] Specifically, in step S5, after the doctor verifies the optimal puncture path and the optimal motion vector, the optimal puncture path and the optimal motion vector are output as navigation suggestions and output to the urinary stone robot (urinary stone robot) to guide or assist the urinary stone robot in operation, thereby improving the accuracy and safety of the surgery.
[0053] Specifically, after outputting the optimal puncture path and optimal motion vector in step S5 to assist the urinary stone robot in performing the surgical operation, the procedure also includes: Obtain surgical outcome data after a urinary calculi procedure performed by a robotic surgical device; Compare the multimodal data and surgical outcome data to determine whether the surgery is complete; if yes, the surgery is confirmed to be complete; if not, the modal data corresponding to the surgical outcome data is determined as the new registered real-time modal data, and the process returns to step S3.
[0054] After the assisted urinary stone removal robot completes the corresponding surgical procedure, it obtains surgical outcome data. This data includes real-time images of the surgical area, real-time CT scan images, real-time probability vectors of stone composition, and tissue feedback after laser treatment.
[0055] Multimodal data before surgery (such as CT scan images and stone composition probability vectors) is compared and analyzed with surgical outcome data obtained after the urinary stone robot performs the corresponding surgical operations to assess the progress and effectiveness of the surgery, thereby determining whether the surgery is complete. For example, image processing technology is used to calculate whether the stone clearance rate is greater than or equal to a preset percentage to quantify the degree of surgical completion. If the stone clearance rate reaches or exceeds the preset percentage, the surgical goal is considered to have been achieved; if not, further processing is required.
[0056] When the judgment result indicates that the surgery is completed (e.g., the stone clearance rate is greater than or equal to the preset percentage), the surgery is confirmed to be completed.
[0057] When the determination result is that the operation is not completed (e.g., the stone clearance rate is less than the preset percentage), the data of the current surgical status is used as the new registered real-time modal data to ensure that subsequent path and action generation is based on the latest and most accurate real-time surgical environment information.
[0058] As shown above, this robot-assisted method for urinary calculi treatment acquires the patient's multimodal data before surgery, uses a Kalman filter algorithm to temporally align the multimodal data, and then spatially registers the patient's real-time video stream data during surgery with the temporally aligned multimodal data using a preset image registration energy function and a preset deformation field model. This yields registered real-time modal data, which is then input into a preset path generation model and a preset motion generation model to generate the optimal puncture path and corresponding optimal motion vector. Finally, the optimal puncture path and optimal motion vector are output to assist in the treatment of urinary calculi. A urinary stone robot performs surgical procedures. By effectively integrating and registering preoperative multimodal data with real-time intraoperative video stream data, and inputting it into a preset path generation model and a preset motion generation model, the robot generates the optimal puncture path and the corresponding optimal motion vector. This assists the urinary stone robot in performing surgical procedures, solving the problems of information fragmentation and "blind puncture" risks in existing urinary stone robots during surgery, which rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. The robot can accurately integrate preoperative and intraoperative information to reduce reliance on the doctor's experience and improve the safety and predictability of urinary stone surgery.
[0059] refer to Figure 2 This application provides a robotic assist device for urinary stone treatment, used to assist a urinary stone treatment robot, comprising: Module 1 is used to acquire multimodal data of patients before surgery; Alignment module 2 is used to perform time alignment on multimodal data using the Kalman filter algorithm to obtain time-aligned multimodal data; The registration module 3 is used to spatially register the real-time video stream data of the patient during surgery with the time-aligned multimodal data through a preset image registration energy function and a preset deformation field model, so as to obtain the registered real-time modal data. The generation module 4 is used to input the registered real-time modal data into the preset path generation model and the preset motion generation model to generate the optimal puncture path and the corresponding optimal motion vector. Output module 5 is used to output the optimal puncture path and optimal motion vector to assist the urinary stone robot in performing surgical operations.
[0060] This robotic-assisted device for urinary stone surgery effectively integrates and registers preoperative multimodal data with real-time intraoperative video stream data, then inputs it into a preset path generation model and a preset motion generation model to generate the optimal puncture path and corresponding optimal motion vector. This assists the urinary stone robot in performing surgical operations, solving the problems of information fragmentation and "blind puncture" risks in existing urinary stone robots during surgery, which rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. It can accurately fuse preoperative and intraoperative information to reduce reliance on the doctor's experience and improve the safety and predictability of urinary stone surgery.
[0061] Specifically, when module 1 is executed, it acquires multimodal data of the patient before surgery using existing technology. The multimodal data includes preoperative CT scan images, MRI images, lithotripsy composition analysis results (calculated based on Fourier transform infrared spectroscopy (FTIR)) and corresponding stone composition probability vectors (the stone composition probability vector represents the probability distribution of different stone components, which can be obtained by inputting the lithotripsy composition analysis results into a pre-set composition probability calculation model to probabilistically evaluate the chemical composition of the stone (such as calcium oxalate, uric acid, magnesium ammonium phosphate, etc.). This composition probability calculation model is a trained machine learning algorithm model) and stone hardness estimation vectors (the stone hardness estimation vector represents the quantitative value or range of stone hardness, which can be obtained by inputting the lithotripsy composition analysis results into a pre-set stone hardness estimation model to estimate the physical hardness of the stone. This stone hardness estimation model is a trained machine learning algorithm model) and other data.
[0062] Specifically, when the alignment module 2 is executed, it uses a precision clock to generate a hardware timestamp and combines a Kalman filter to make the optimal estimate of the offset between the master and slave clocks, providing a unified time reference with an error of less than 1ms for multimodal data, so as to perform time alignment on the multimodal data and obtain time-aligned multimodal data, thereby eliminating data inconsistency caused by differences in acquisition time and ensuring that all preoperative data are consistent in the time dimension.
[0063] The state equation for the Kalman filter is as follows: ; in, Let k be the state vector (i.e., clock offset). Let k be the state vector at time k-1; Here is the state transition matrix. , The sampling period; This is process noise, reflecting the rate of change of the clock offset, and can be set according to actual needs.
[0064] The observation equation for Kalman filtering is as follows: ; in, Let k be the observation vector (i.e., the actual time data). For the observation matrix, ; To observe noise, settings can be adjusted based on the calculation error of the time data.
[0065] The state equation and observation equation are iterated using Kalman filtering: ; ; ; in, The Kalman gain at time k; Let be the error covariance matrix at time k-1; Let be the error covariance matrix at time k; The predicted clock offset at time k; This is the predicted clock offset at time k-1.
[0066] In summary, a unified time base is calculated to align the time of multimodal data, resulting in time-aligned multimodal data.
[0067] Specifically, when the registration module 3 spatially registers the real-time video stream data of the patient during surgery with the time-aligned multimodal data using a preset image registration energy function and a preset deformation field model to obtain the registered real-time modal data, it executes the following: Real-time acquisition of video stream data of patients during surgery; The real-time video stream data is compensated by a preset deformation field model to obtain compensated real-time video stream data. Using a preset image registration energy function, spatial registration is performed on the compensated real-time video stream data and the time-aligned multimodal data to obtain the registered real-time modal data.
[0068] During the robot-assisted urological stone removal process, registration module 3 continuously acquires video image sequences of the patient's surgical area using an endoscope or other real-time imaging equipment. These sequences are then combined to form a continuous sequence of image frames, resulting in real-time video stream data. This video stream data reflects the dynamic changes in the surgical environment and forms the basis for real-time control and operation.
[0069] Due to respiratory movements during surgery, the urinary system experiences rhythmic displacement, which can cause dynamic deformation and interfere with subsequent registration. Therefore, a deformation field model is designed to address this issue, ensuring that the compensated real-time video stream data more accurately reflects the true anatomical structure of the surgical area, thus reducing the interference of dynamic deformation on subsequent registration. The pre-designed deformation field model is as follows: ; Among them, among them, For the compensated surgical time (i.e., in the compensated real-time video stream data) The spatial coordinates of point m after the deformation field has been applied; For the point during surgery (i.e., in the real-time video stream data) The spatial coordinates of point m, that is, the spatial coordinates of point m after it has been subjected to the deformation field. This indicates that static deformation due to respiratory movements is not considered. , For points before surgery (i.e., in time-aligned multimodal data) spatial coordinates, For point m (to point The displacement of ) The spatial distribution of the effects of respiratory motion can be learned from 4D-CT (four-dimensional CT scan images) through principal component analysis (PCA); It is a time function of the respiratory signal and can be obtained by real-time observation of thoracic impedance, vPPG (virtual portal vein pressure gradient), or external markers.
[0070] Using a preset image registration energy function, spatial registration is performed on the compensated real-time video stream data and the time-aligned multimodal data. The optimal registration transformation is found by optimizing the image registration energy function to obtain the registered real-time modal data.
[0071] Specifically, the steps for constructing the preset image registration energy function include: Based on the similarity of the images, an initial image registration energy function is constructed; By introducing a deformation field regularization term, the initial image registration energy function is optimized to obtain a preset image registration energy function.
[0072] When constructing the preset image registration energy function, an initial image registration energy function is obtained based on the image similarity (such as the similarity between corresponding CT scan images) in real-time video stream data and multimodal data. Specifically, the initial image registration energy function is as follows: ; in, For optimal deformation field; Register the energy function for the initial image; Let this be the initial objective function; This is a similarity metric used to calculate similarity. Preoperative 3D grayscale CT images in DICOM format obtained by a CT scanner are used to provide high-resolution anatomical structures. for Image midpoint Position coordinates; This method allows for the real-time acquisition of two-dimensional / three-dimensional grayscale images of anatomical structures via ultrasound probes or endoscopic cameras during surgery. Intraoperative ultrasound / endoscopic images The coordinates of the midpoint m; , It is a deep feature extractor, which extracts high-level semantic features from images. The square of the L2 norm, The difference between two feature vectors is measured by calculating the Euclidean distance between the corresponding feature vectors of the CT scan image and the ultrasound / endoscopic image; the initial image registration energy function. This indicates that the initial objective function is... When the minimum value is obtained, and The position makes the optimal deformation field To achieve the optimal state, i.e., when the image is most similar, determine the optimal deformation field at this point. It can achieve registration.
[0073] The purpose of introducing a deformation field regularization term is to constrain the smoothness and physical rationality of the deformation field, avoiding irregular or discontinuous jumps, thereby improving the robustness and accuracy of the registration results. Specifically, the deformation field regularization term is constrained using a viscous fluid model, which is defined as follows: ; in, For deformation field regularization term; It is a Laplace operator, which can calculate the curvature of the deformation field using the finite difference method; Let i be the i-th displacement component. , This represents the first component of the displacement, that is, the displacement in the horizontal direction. This is the second displacement component, that is, the displacement in the vertical direction. The third displacement component is the displacement in the vertical direction, where x is the abscissa of point m, y is the ordinate of point m, and z is the ordinate of point m.
[0074] In summary, the preset image registration energy function is as follows: .
[0075] Therefore, a deformation field regularization term is introduced into the initial image registration energy function to obtain a preset image registration energy function. This imposes constraints on the smoothness and continuity of the deformation field, ensuring that during the optimization process, not only image similarity but also the physical rationality of the deformation field must be considered. This combined approach allows the registration process to maintain image matching accuracy while effectively suppressing the ill-conditioned behavior of the deformation field, thereby obtaining more stable and reliable registration results.
[0076] Specifically, when the generation module 4 inputs the registered real-time modal data into the preset path generation model and the preset motion generation model to generate the optimal puncture path and the corresponding optimal motion vector, it executes the following: The registered real-time modal data is input into the preset feature enhancement model to obtain the stone feature enhancement image; The probability vector of stone composition and the estimation vector of stone hardness are extracted from the registered real-time modal data. The enhanced image of the stone features, the probability vector of the stone composition, and the estimation vector of the stone hardness are input into the preset path generation model and the preset action generation model to generate the optimal puncture path and the corresponding optimal action vector.
[0077] During execution, generation module 4 inputs the registered real-time modal data into a preset feature enhancement model to highlight the features of the stone region, suppress noise, and enhance the contrast or texture information of the image, resulting in a stone feature-enhanced image. This provides clearer and more discriminative visual information for subsequent path and action generation. Feature enhancement makes key information such as the stone's boundaries and internal structure more apparent, helping the preset path generation model and action generation model to more accurately identify and locate the stone. The feature enhancement model is a deep learning model, such as a convolutional neural network (CNN) or a generative adversarial network (GAN). Deep learning models are existing technologies and will not be detailed here.
[0078] From the registered real-time modal data, the probability vector of stone composition and the estimation vector of stone hardness are extracted. These vectors provide information on the intrinsic properties of the stone, which is crucial for selecting appropriate treatment parameters such as laser power and frequency of the urinary stone robot.
[0079] Specifically, before inputting the registered real-time modal data into the preset feature enhancement model to obtain the stone feature enhancement image, the generation module 4 also performs the following: Using a pre-trained encoder, the registered real-time modal data and the pose information of the urinary stone robot are mapped into a common space, so that multiple data at the same surgical moment are located in the same common space.
[0080] When generation module 4 is executed, before inputting the registered real-time modal data into the preset feature enhancement model, a pre-trained encoder can be used to map the registered real-time modal data and the pose information of the urinary stone robot into a 512-dimensional common space. For example, ResNet-50 (or other CNNs) can be used to extract image features from the registered real-time modal data, and then projected to 512 dimensions through a fully connected layer; 1D-CNN can be used to extract features from the lithotripsy component analysis results in the registered real-time modal data, and then projected to 512 dimensions; MLP (Multilayer Perceptron) can be used to extract the pose features of the urinary stone robot and project them to 512 dimensions. By mapping high-dimensional, heterogeneous real-time modal data and the pose information of the urinary stone robot into the common space through a pre-trained encoder, the purpose is to narrow the representation distance between different modalities of the same sample, making multiple data at the same surgical moment close in the common space, thereby effectively reducing the dimensionality of the data, removing redundant information, and bridging the semantic gap between different data sources. In the public space, these data are uniformly represented as low-dimensional, semantically consistent feature vectors, enabling subsequent path generation and action generation models to extract key features more efficiently and accurately.
[0081] Specifically, when the generation module 4 inputs the feature-enhanced image, the stone composition probability vector, and the stone hardness estimation vector into the preset path generation model and the preset action generation model to generate the optimal puncture path and the corresponding optimal action vector, it performs the following: The feature-enhanced image is input into a preset path generation model to generate the optimal puncture path; The optimal puncture path, feature-enhanced image, stone composition probability vector, and stone hardness estimation vector are input into the preset action generation model to generate the optimal action vector corresponding to the optimal puncture path.
[0082] During execution, the generation module 4 inputs the enhanced feature image into a preset path generation model. Based on the spatial information and pathological features provided by the enhanced image, the model plans an optimal puncture path from the urinary stone robot's operating point to the urinary stone, avoiding critical tissues, blood vessels, or nerves, and ensuring the safety and effectiveness of the puncture procedure. After generating the optimal puncture path, it must be reviewed by a physician before proceeding to the next step. The path generation model is a model built by combining existing path search algorithms (such as Dijkstra's algorithm, A* algorithm, or RRT fast random tree algorithm) and existing deep learning algorithms (such as neural network algorithms). It can generate a safe, efficient, and executable virtual puncture channel from the skin entry point to the target stone based on both static and dynamic anatomical environments.
[0083] The optimal puncture path, along with the enhanced image, stone composition probability vector, and stone hardness estimation vector, are input into a pre-defined motion generation model. Based on this comprehensive information, the model generates an optimal motion vector that highly matches the optimal puncture path. This optimal motion vector represents a series of precise instructions required for the urinary stone robot to perform the puncture operation, including optimal puncture depth, optimal puncture angle, optimal laser power, optimal frequency, and optimal perfusion pressure. The motion generation model, constructed using existing reinforcement learning algorithms (such as Q-learning), transforms the planned path (optimal puncture path) into specific, precise motion commands and energy parameters for the end effector (puncture needle / laser fiber) of the urinary stone robot at each moment.
[0084] Specifically, when the output module 5 is executed, after the doctor verifies the optimal puncture path and the optimal motion vector, it outputs the optimal puncture path and the optimal motion vector as navigation suggestions to the urinary stone robot (urinary stone robot) to guide or assist the urinary stone robot in its operation, thereby improving the accuracy and safety of the surgery.
[0085] Specifically, after outputting the optimal puncture path and optimal motion vector to assist the urinary stone robot in performing surgical operations, output module 5 also performs: Obtain surgical outcome data after a urinary calculi procedure performed by a robotic surgical device; The multimodal data and surgical outcome data are compared to determine whether the surgery is completed. If yes, the surgery is confirmed to be completed. If no, the modal data corresponding to the surgical outcome data is determined as the new registered real-time modal data, and the registration module 3 is controlled to execute the corresponding steps.
[0086] After the assisted urinary stone removal robot completes the corresponding surgical procedure, it obtains surgical outcome data. This data includes real-time images of the surgical area, real-time CT scan images, real-time probability vectors of stone composition, and tissue feedback after laser treatment.
[0087] Output module 5 compares and analyzes preoperative multimodal data (such as CT scan images, stone composition probability vectors, etc.) with surgical result data obtained after the urinary stone robot performs the corresponding surgical operation to evaluate the progress and effect of the surgery, thereby determining whether the surgery is completed. For example, image processing technology is used to calculate whether the stone clearance rate is greater than or equal to a preset percentage to quantify the completion of the surgery. If the stone clearance rate reaches or exceeds the preset percentage, the surgical goal is considered to have been achieved; if not, further processing is required.
[0088] When the judgment result indicates that the surgery is completed (e.g., the stone clearance rate is greater than or equal to the preset percentage), the surgery is confirmed to be completed.
[0089] When the determination result is that the operation is not completed (e.g., the stone clearance rate is less than the preset percentage), the data of the current surgical state is used as the new registered real-time modal data, and the registration module 3 is controlled to execute the corresponding steps to ensure that the subsequent path and action generation is based on the latest and most accurate real-time surgical environment information.
[0090] As described above, this robotic-assisted device for urinary calculi acquisition acquires the patient's multimodal data before surgery, uses a Kalman filter algorithm to perform time alignment on the multimodal data, and obtains time-aligned multimodal data. Then, using a preset image registration energy function and a preset deformation field model, it spatially registers the patient's real-time video stream data during surgery with the time-aligned multimodal data, obtaining registered real-time modal data. This registered real-time modal data is then input into a preset path generation model and a preset motion generation model to generate the optimal puncture path and corresponding optimal motion vector. Finally, it outputs the optimal puncture path and optimal motion vector to assist in the procedure. A urinary stone robot performs surgical procedures. By effectively integrating and registering preoperative multimodal data with real-time intraoperative video stream data, and inputting it into a preset path generation model and a preset motion generation model, the robot generates the optimal puncture path and the corresponding optimal motion vector. This assists the urinary stone robot in performing surgical procedures, solving the problems of information fragmentation and "blind puncture" risks in existing urinary stone robots during surgery, which rely on the doctor's personal experience, resulting in low treatment efficiency and excessive risks. The robot can accurately integrate preoperative and intraoperative information to reduce reliance on the doctor's experience and improve the safety and predictability of urinary stone surgery.
[0091] Please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device includes a processor 301 and a memory 302. The processor 301 and the memory 302 are interconnected and communicate with each other via a communication bus 303 and / or other connection mechanisms (not shown). The memory 302 stores a computer program executable by the processor 301. When the electronic device is running, the processor 301 executes the computer program to perform the urinary stone robot-assisted method in any optional implementation of the above embodiments, to achieve the following functions: acquiring multimodal data of the patient before surgery; using a Kalman filter algorithm to perform time alignment on the multimodal data to obtain time-aligned multimodal data; spatially registering the real-time video stream data of the patient during surgery and the time-aligned multimodal data using a preset image registration energy function and a preset deformation field model to obtain registered real-time modal data; inputting the registered real-time modal data into a preset path generation model and a preset motion generation model to generate an optimal puncture path and a corresponding optimal motion vector; and outputting the optimal puncture path and optimal motion vector to assist the urinary stone robot in performing surgical operations.
[0092] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the urinary stone robot-assisted method in any optional implementation of the above embodiments to achieve the following functions: acquiring the patient's multimodal data before surgery; using a Kalman filter algorithm to perform time alignment on the multimodal data to obtain time-aligned multimodal data; spatially registering the patient's real-time video stream data during surgery and the time-aligned multimodal data using a preset image registration energy function and a preset deformation field model to obtain registered real-time modal data; inputting the registered real-time modal data into a preset path generation model and a preset motion generation model to generate an optimal puncture path and a corresponding optimal motion vector; and outputting the optimal puncture path and the optimal motion vector to assist the urinary stone robot in performing surgical operations. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0093] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0094] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0095] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0096] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.
[0097] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A robot-assisted method for treating urinary stones, used to assist a robot for treating urinary stones, characterized in that, Including the following steps: S1, acquire multimodal data of the patient before surgery; S2, The Kalman filter algorithm is used to perform time alignment on the multimodal data to obtain time-aligned multimodal data; S3, using a preset image registration energy function and a preset deformation field model, spatially register the real-time video stream data of the patient during surgery with the time-aligned multimodal data to obtain the registered real-time modal data; S4, input the registered real-time modal data into the preset path generation model and the preset action generation model to generate the optimal puncture path and the corresponding optimal action vector. S5, output the optimal puncture path and the optimal motion vector to assist the urinary stone robot in performing surgical operations.
2. The robot-assisted method for urinary calculi treatment according to claim 1, characterized in that, Using a preset image registration energy function and a preset deformation field model, the real-time video stream data of the patient during surgery and the time-aligned multimodal data are spatially registered to obtain registered real-time modal data, including: Real-time acquisition of the patient's video stream data during surgery; The real-time video stream data is compensated using a preset deformation field model to obtain compensated real-time video stream data. Using a preset image registration energy function, spatial registration is performed on the compensated real-time video stream data and the time-aligned multimodal data to obtain registered real-time modal data.
3. The robot-assisted method for urinary calculi treatment according to claim 2, characterized in that, The steps for constructing the preset image registration energy function include: Based on the similarity of the images, an initial image registration energy function is constructed; A deformation field regularization term is introduced to optimize the initial image registration energy function, thereby obtaining the preset image registration energy function.
4. The robot-assisted method for urinary calculi treatment according to claim 1, characterized in that, The registered real-time modal data is input into a preset path generation model and a preset motion generation model to generate the optimal puncture path and the corresponding optimal motion vector, including: The registered real-time modal data is input into a preset feature enhancement model to obtain a stone feature enhancement image; The stone composition probability vector and stone hardness estimation vector are extracted from the registered real-time modal data. The enhanced image of the stone features, the probability vector of the stone composition, and the estimation vector of the stone hardness are input into a preset path generation model and a preset action generation model to generate the optimal puncture path and the corresponding optimal action vector.
5. The robot-assisted method for urinary calculi treatment according to claim 4, characterized in that, Before inputting the registered real-time modal data into a preset feature enhancement model to obtain the stone feature enhancement image, the process also includes: Using a pre-trained encoder, the registered real-time modal data and the pose information of the urinary stone robot are mapped into a common space, so that multiple data at the same surgical moment are located in the same common space.
6. The robot-assisted method for urinary calculi treatment according to claim 5, characterized in that, The enhanced image of the stone features, the probability vector of the stone composition, and the estimation vector of the stone hardness are input into a preset path generation model and a preset action generation model to generate the optimal puncture path and the corresponding optimal action vector, including: The enhanced image of the stone features is input into a preset path generation model to generate the optimal puncture path; The optimal puncture path, the enhanced image of the stone features, the probability vector of the stone composition, and the estimation vector of the stone hardness are input into a preset action generation model to generate the optimal action vector corresponding to the optimal puncture path.
7. The robot-assisted method for urinary calculi treatment according to claim 6, characterized in that, After outputting the optimal puncture path and the optimal motion vector to assist the urinary stone robot in performing surgical operations, the procedure also includes: Obtain surgical outcome data after a urinary calculi procedure performed by a robotic surgical device; The multimodal data and the surgical result data are compared to determine whether the surgery is completed. If yes, the surgery is confirmed to be completed. If no, the modal data corresponding to the surgical result data is determined as the new registered real-time modal data, and the process returns to step S3.
8. A urinary stone robot-assisted device for assisting a urinary stone robot, characterized in that, include: The acquisition module is used to acquire multimodal data of patients before surgery; The alignment module is used to perform time alignment on the multimodal data using a Kalman filter algorithm to obtain time-aligned multimodal data. The registration module is used to spatially register the real-time video stream data of the patient during surgery with the time-aligned multimodal data using a preset image registration energy function and a preset deformation field model, so as to obtain the registered real-time modal data. The generation module is used to input the registered real-time modal data into a preset path generation model and a preset action generation model to generate the optimal puncture path and the corresponding optimal action vector. The output module is used to output the optimal puncture path and the optimal motion vector to assist the urinary stone robot in performing surgical operations.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program executable by the processor, which, when executing the computer program, performs the steps of the robot-assisted method for urinary stones as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it performs the steps of the robot-assisted method for urinary stones as described in any one of claims 1-7.