Machine learning network generated by a medical robot according to configurable modules

Through generative adversarial networks and machine learning technology, the optimal configuration of modular robot systems is automatically determined, which solves the design problems of complex tasks and heterogeneous module combinations in the existing technology, and achieves efficient and safe complex environment operations.

CN114945925BActive Publication Date: 2025-07-18SIEMENS HEALTHINEERS AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080094041.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-24
Publication Date
2025-07-18
Estimated Expiration
2040-01-24

AI Technical Summary

Technical Problem

When existing modular robot systems combine complex tasks and heterogeneous modules, it is difficult to determine the optimal parameters and arrangement, resulting in high design complexity and difficulty in computing, and cannot effectively deal with complex morphology and programming assembly requirements.

Method used

Generative adversarial networks (GAN) and machine learning technology are used to automatically determine the optimal configuration of a modular robot system by training neural networks, and use anatomical structure models to constrain performance and operations to generate robot configurations that meet task-specific specifications.

Benefits of technology

The optimal configuration of modular robotic systems is achieved in complex environments, improving design efficiency and safety, and being able to effectively perform complex medical tasks such as ultrasound imaging and spinal correction surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114945925B_ABST
    Figure CN114945925B_ABST
Patent Text Reader

Abstract

A generative adversarial network (GAN) (21, 24) or any other generative modeling technique is used to learn (12) how to generate (68) an optimal robot system given performance, operation, safety, or any other specification. For example, the specification can be modeled (65) relative to an anatomical structure to confirm satisfaction of anatomical structure-based constraints or another task-specific constraint. A machine learning system (e.g., a neural network) is trained (12) to translate a given specification into a robot configuration. The network can translate a task-specific specification into one or more configurations of robot modules in the robot system. A user can enter (67) a change to the performance so that the network estimates (62) an appropriate configuration. The configuration can be translated (64) into an estimated performance by another machine learning system (e.g., a neural network), thereby allowing the operation to be modeled (65) relative to an anatomical structure (such as a medical imaging-based anatomical structure). Configurations that satisfy the constraints from the modeling (65) can be assembled (69) and used.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This embodiment relates to configuring a modular robot system. Modular robots are optimized or configured manually or iteratively by a human designer for a specific task. The configuration is parameterized using discrete and continuous variables, so only a few settings of the parameters for the configuration are often evaluated during the design process.

[0002] Self-reconfiguring robots (e.g., robotic units, electronic device modules, and / or software modules in a basic lattice structure) can metamorph into different configurations. Although heterogeneous modular systems have been developed, each module still remains too simple, having two or three degrees of freedom for assembling into a chain-like or lattice-like structure. The types of modules are often limited to one or two types. In a chain-like structure, a single module implementing a rotary joint is repeatedly combined. In a lattice-like structure, the modules can form a polyhedral structure. Often, a random process is followed to achieve the desired configuration. However, this approach fails to generalize to complex morphologies for performing specific tasks that are far from the capabilities of a single module, or fails to respond to programmable assembly requirements. For more complex modular components, such as a modular component with five links having five configurable lengths and three or more types of such modules, even the complexity of determining the optimal parameters and arrangements based on a specific two-dimensional workspace given specific requirements makes brute force calculation of the configuration difficult. Summary of the Invention

[0003] As an introduction, the preferred embodiments described below include methods, computer-readable media, and systems for machine training and application of machine learning models to configure modular robotic systems, such as but not limited to robots for medical imaging or applications for specific types of medical applications or anatomy. Generative adversarial networks (GANs) or any other generative modeling techniques are used to learn how to generate an optimal robotic system given performance, operation, safety, or any other specifications. For example, specifications can be modeled relative to anatomy to confirm satisfaction of anatomy-based constraints or other task-specific constraints. A machine learning system (e.g., a neural network) is trained to translate given specifications into robotic configurations. The network can translate task-specific specifications into one or more configurations of robotic modules in a robotic system. A user can enter changes to performance so that the network estimates appropriate configurations. The configurations can be translated into estimated performance by another machine learning system (e.g., a neural network), allowing operation to be modeled relative to anatomy (e.g., anatomy based on medical imaging). Configurations that satisfy the constraints from the modeling can be assembled and used.

[0004] In a first aspect, a method for generating a medical robot from configurable modules is provided. A first ability of the input medical robot, such as a task-based desired ability, is input. A machine learning encoder projects the first ability onto a latent space vector that defines the configuration of the configurable modules. A machine learning generator estimates a second ability based on the latent space vector as input to the machine learning generator. The operation of the configured medical robot is modeled based on the second ability relative to an anatomy model. The configuration of the configurable modules is set based on the result of the modeling and the latent space vector.

[0005] In one embodiment, force, motion, compliance, workspace, payload, and / or joint position are input as abilities. Abilities for robotic structures or modules, circuits, artificial intelligence (AI), enclosures, and / or human-machine interfaces can be input. The abilities are projected onto a latent space vector. For example, the latent space vector includes values of parameters for the type of configurable module, connections between configurable modules, and adjustable aspects of the configurable modules. The latent space vector can be the shape of the configured robot (i.e., the assembly of components).

[0006] In another embodiment, the machine learning generator is a neural network. The neural network or other generator may have been trained as a generative adversarial network.

[0007] According to one embodiment, the generated ability is a three-dimensional vector. For example, force or compliance is provided in three dimensions as the output of the generator.

[0008] In other embodiments, the processor receives user input that changes one of the estimates of one of the second capabilities. The repetition of the projection using the estimates of the second capabilities including the changed estimate is repeated. The repetition of the generation of another latent space vector projected by the repetition of the projection is repeated. The repetition of the modeling of the third capability obtained from the repetition of the generation is repeated. The user input may be restricted to a change in only one of the estimates of the second capabilities. The user input is used to change one estimate without a change in the configuration. In one example, the change is in the direction of the estimate of the second capability.

[0009] In one embodiment, a trained computational model from image data is used in the modeling.

[0010] In a second aspect, a method for machine training to configure a robot according to component modules is provided. Various configurations of the robot from component models and training data of the performance of the robot are provided. A generative adversarial network is machine trained to estimate performance from the input of the configuration. The generative adversarial network includes a discriminator configured to map the performance to a real or fake classification, wherein the machine training includes constraints based on modeling the performance with respect to an anatomical structure. The machine-trained generator of the generative adversarial network is stored. In a further embodiment, an encoder may be machine trained to project the performance to the configuration.

[0011] In one embodiment, the machine training includes constraints on modeling the configuration with respect to an anatomical structure. The modeling provides an adjusted value of the performance. The encoder is machine trained to project the performance and / or the adjusted performance to the configuration.

[0012] In another embodiment, modeling the configuration with respect to an anatomical structure includes modeling using an anatomical structure model from medical imaging.

[0013] In a third aspect, a method for generating a medical robot according to configurable modules is provided. A machine learning encoder projects from a user-defined workspace, joint space, force space, compliance space, and / or kinematic space to the configuration and parameter space of the configurable module. The projection may be from a space for a robot structure or module, circuit, artificial intelligence (AI), packaging, and / or human-machine interface. The machine learning encoder is trained based on a generative adversarial network that estimates values in the workspace, joint space, force space, compliance space, kinematic space, and / or other spaces from the configuration and parameter space. According to the projection, the configuration of the configurable module is determined for the user-defined workspace, joint space, force space, compliance space, and / or kinematic space.

[0014] In another embodiment, a machine learning generator of a generative adversarial network generates estimates of a workspace, a joint space, a force space, a compliance space, and / or a kinematic space from values of a configuration and a parameter space of a configurable module. User edits to the estimates are received. The projection is repeated from the edited estimates.

[0015] In yet another embodiment, a user-defined workspace, joint space, force space, compliance space, and / or kinematic space is constrained based on modeling an interaction with an anatomical structure.

[0016] The invention is defined by the following claims, and nothing in this section should be taken as a limitation on those claims. Features of one type of claim (e.g., method or system) may be used in another type of claim. Further aspects and advantages of the invention are discussed in connection with the preferred embodiments below, and these aspects and advantages may later be claimed independently or in combination. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The components and the figures are not necessarily to scale; instead, emphasis is placed upon illustrating the principles of the invention. In addition, in the figures, like reference numerals designate corresponding parts throughout the different views.

[0018] Figure 1 is a flowchart of an embodiment of a method for training for a medical robot configuration;

[0019] Figure 2 illustrates an example arrangement for machine training for a robot configuration;

[0020] Figure 3 shows three different modules configured to form a robot;

[0021] Figure 4 and 5 shows other example robots configured according to various module combinations;

[0022] Figure 6 is a flowchart of an embodiment of a method for generating a medical robot according to configurable modules; and

[0023] Figure 7 is a block diagram of an embodiment of a system for configuring a robot system according to modular components. DETAILED DESCRIPTION

[0024] Unlike conventional modular robots, the configured robots are intended to perform specific tasks in a medical environment. Modular heterogeneous components can be combined in a hybrid system to perform complex movements and tasks in complex environments that are difficult to model and predict. Examples of such tasks include manipulating an ultrasound transducer to obtain images of anatomical structures such as the heart or prostate, guiding a catheter device to a desired target using online feedback from images and sensing, holding and positioning a medical tool on a patient's skin (i.e., a tool holder mounted to the patient), or guiding a robotic system for a spinal correction procedure. In these applications, it may be important not only to determine the qualities of the robot and design for the robot's qualities, but also to design the robot to effectively and safely interact with these complex environments.

[0025] Generative adversarial networks (GANs) are used to generate de novo a medical robotic system from heterogeneous reconfigurable modules. A specific high-level description of the tasks and / or specifications that the module must satisfy can be input to determine the configuration without manual design or with minimal manual design. The optimal configuration and parameters of each module are determined using one or more machine learning models. Training using GANs allows the (one or more) user descriptions related to the task or specification to be back-projected into the latent space of the configuration and parameter space of the constitutive modules. The entire process of providing the task description and specification to the generation of the configuration and parameter space of the constitutive modules can be automated.

[0026] The use of GANs allows the user to interactively edit along the latent space of the model by editing the task or specification. Multiple specifications can be provided to determine the desired configuration. In addition to the Mx3D Cartesian channels (where M is the number of components in the current configuration of the system), the model also includes spaces such as workspace, joint space, force space, compliance, etc. as part of the specification.

[0027] The model can include parameters associated with the electronics that control the robot, the underlying logic or AI software, the enclosure, the human-machine interface, the robotic structure or modules, and / or other elements of the robotic system. For example, electronic modules, algorithm modules, and / or mechanical modules are used. A number of controllers, chip sets, memories, circuits, and / or software modules can be provided at the sensor level, actuators, logic / AI, or other locations.

[0028] Data-driven learning can be used to generate a virtual environment. For example, medical scan data is used to obtain a parameterized computational model of a specific anatomical structure. . Having such a model allows determination based on the generator via GAN How the design of a given latent vector to be mapped is to be executed under the selection conditions in a virtual environment. Additionally, machine learning methods (e.g., deep reinforcement learning) can be deployed to further refine the design.

[0029] Figure 1 An embodiment of a flowchart of a method for machine training to configure a robot according to component modules is shown. The goal is to train a neural network to generate a configuration of a modular robot according to high-level descriptions (such as category, workspace requirements, payload capacity, etc.). The encoder is trained, as part of training a GAN, for inverse transformation - from the configuration of the robot to a performance description. Anatomical structure modeling can be used to constrain the performance and thus the configuration used.

[0030] Figure 2 An example of different machine training modules and their relationships to each other is illustrated. The generator 21 and the discriminator 24 constitute a GAN that is machine-trained with a projection encoder 26. The anatomical structure model 23 can be machine-trained, pre-trained, or can be a model simulation based on physics or other biometrics without machine training. Any one or more of the models can be pre-trained and used in the training of one or more of the other models. Alternatively, joint training or end-to-end training is used to train all or multiple of the models as part of the same optimization. Figure 2 An illustration of the interaction between different models during training and once trained is shown.

[0031] Figure 1 The method is implemented by a processor such as a computer or a server. For example, using Figure 7 system 70 of

[0032] The method is executed in the order shown or in other orders. For example, action 14 is executed before action 14, as part of action 14, or after action 14.

[0033] Additional, different, or fewer actions can be provided. For example, actions for separately training the GAN, the generator of the GAN, the discriminator of the GAN, the encoder, and / or the anatomical structure model or the robot-to-anatomical structure model interaction are provided.

[0034] In action 10, training data is collected. To provide training data, many samples or examples are collected in one or more memories.

[0035] The training data includes various configurations of the robot from the component models and the performance of the robot. The training data set includes those with a target O ={( x 1 , y1 , s 1 ),…,( x n , y 1 , s n )} example D ={( c 1 , θ 1 ),…,( c n , θ n )}. The input is a variable that encodes the category (type of sub-module) c , and the parameters of the sub-module θ , together with the tuples of the interconnections between the included modules (Example D). These tuples form the latent space vector 20. The latent space vector 20 provides configurations such as which modules are used, the interconnections or arrangements of the used modules, and any adjustable or settable parameters of each module (e.g., the length of an arm or link, or the range of motion). In one embodiment, the latent space vector 20 is the shape of the robot given by the selected modules, the interconnections between the modules, and / or the set values of the parameters of each module. In other embodiments, the latent space vector 20 includes dynamic information such as changes or shapes over time in one or more variables (e.g., the configuration includes changes over time in any aspect of the configuration). Step increments in position and / or velocity may be included in the configuration to account for changes in shape over time.

[0036] The modules to be incorporated into the robot can belong to any type or variant. Any number of modules of each type are available. Any number of module types are available. Different robot modules can be combined or interconnected to form a robot or a robot system for a given task. These modules are reconfigurable, so different groups of modules can be used to form different medical robot systems.

[0037] Figure 3An example is shown. Three types of modules - including a z-θ module 30, an x-y stage module 32, and an instrument holder module 34 - are combined to form a robotic system for image-guided spinal surgery. These heterogeneous components can be selected and interconnected to form a robot. The z-θ module 30 is a combined prismatic and rotational joint made using a splined screw mechanism. The rotational rate, threads, and / or length of the threaded part can be configurable parameters. The x-y stage module 32 includes a five-bar mechanism with five link length parameters. Other parameters can be provided, such as the link moment of inertia that determines the stiffness of the mechanism. There may not be a one-to-one mapping between these parameters (such as the moment of inertia) and the latent variables because there are several ways to increase the moment of inertia, such as I-channels or even custom-shaped channels. Machine learning is used to find the optimal mapping given the training data. The instrument holder module 34 includes an active clamping or holding component. The parameters of the instrument holder module 34 can include length, inner diameter, outer diameter, and / or connection type.

[0038] Other types of modules can be provided. For example, an extension arm as shown in Figure 4 is used. As another example, a remote center of motion module as shown in Figure 5 is used. The remote center of motion module is a spatial five-bar or link arrangement for controlling the angulation of a device. In Figure 5 , one or more links are used to rest on the patient.

[0039] These modules have any degrees of freedom for rotation and / or translation, such as the two degrees of freedom provided for the x-y stage module 32, the z-θ module 30, or the remote center of motion module. These modules are either active or passive. Example passive components include sensors or interfaces, such as an interface with a medical device (such as a laparoscope or an ultrasound transducer). Any number of different modules can be provided or used in the design and assembly into a robotic system. The same components or the same type of components can be assembled into different robotic systems.

[0040] Similar modularity applies to electronics, AI architectures and software, and the packaging and other components necessary for building a robotic system. The same electronics, artificial intelligence architectures, software modules, packaging, human-machine interfaces, and / or other components of a robotic system can be assembled into different robotic systems. The examples herein are for robotic structures, but alternatively or additionally, can be for machine-learned configurations of specifications for other types of systems based on robotic systems (e.g., electronics, software, or human-machine interfaces).

[0041] In the case ofFigure 3 In the robotic guidance of spinal surgery, the active module is designed to provide higher force capabilities by increasing the moment of inertia of five rods. The image-driven property specification results in the passive module having imagable components, such as those that may appear in ultrasound or x-ray imaging.

[0042] Figure 4 An example configuration of the interconnection from three modules is shown. The three modules are an arm, an x-y platform, and an instrument holder that provides movement along a curved surface. Figure 4 The robot is a transrectal ultrasound robot. The active module is reconfigured to provide force capabilities suitable for the task by adjusting and / or limiting the drive torque. The selective compliance of the robot is also reconfigured to limit compliance along a user-defined axis.

[0043] Figure 5 An example configuration of the interconnection from two modules is shown. The two modules are a remote center motion platform and an electromechanical laparoscope holder. Figure 5 The robot is a "tool" holder robot mounted to the patient. The active module is reconfigured to provide force capabilities suitable for the task by adjusting and / or limiting the drive torque.

[0044] The training data includes many examples of different configurations. Parameters, the category of modules used, and the values of the interconnections are provided for different configurations. For example, multiple samples of robots for different tasks are provided (e.g., Figures 3 - 5 ). For a given task, different configurations can be provided by changing or using different values of any of the various parameters (e.g., module settings, number of modules, module type, and / or module interconnection). The training data is generated from simulations and / or assemblies.

[0045] The training data also includes the performance of each configuration. Physics-based configurations result in various capabilities of the robot. The capabilities of a given robot are provided for machine training. The performance can be a goal, which is the ND space 22 of the system. The ND space 22 characterizes the performance of the robot as configured, such as operation parameterization, capabilities, safety bounds, and / or other performance indicators. Example performances or capabilities include: force, workspace, joint space, force space, compliance, or other information. For example, the motion limits, space limits, amount of force, compliance level, kinematics, or other operating characteristics of the robot are provided. Any specification can be used, such as specifications for mechanical, software, electronics or circuits, packaging, and / or human-machine interface specifications.

[0046] The ND space 22 can have multiple 3D channels. Each channel represents a 3D model of a sub-module component that is augmented with other channels as desired by the user, including but not limited to a workspace, a tensor representing force application, load capacity, etc. The discriminator maps the ND space 22 to a binary output indicating true or false.

[0047] The parameters of the ND space 22 can be the specifications created by the desired task. Instead of designing based on testing different module combinations, the design is automatically completed by providing the desired performance, such that the machine learning system determines the configuration based on this desired performance. The ability of the user to describe using a high-level description (such as a cylindrical tool holder having a specified sensitivity to force measurements along a particular direction) can be used for the design. Once the user describes the cylindrical tool, the projection encoder 26 is to be trained to project the performance into the latent space of the system. There may be couplings and / or dependencies between the various ND channels output by the generator 21. The GAN learns to abstract this coupling from the user, thereby allowing them to only need to specify the high-level description.

[0048] In action 12, the processor performs machine training. Figure 2 The arrangement of... is for machine training. The generator 21 and the projection encoder 26 are trained sequentially or jointly. The generator 21, the discriminator 24, and the projection encoder 26 are neural networks, such as deep convolutional neural networks. Fully connected, dense, or other neural networks can be used. The same or different network architectures are provided for each of the generator 21, the discriminator 24, and the projection encoder 26. Any number of layers, features, nodes, connections, convolutions, or other architectural aspects of the neural network can be provided.

[0049] In an alternative embodiment of the GAN, other neural networks are machine-trained, such as DenseNet, convolutional neural networks, fully connected neural networks, or three or more layers forming a neural network. Other generative networks without adversarial training (i.e., without the discriminator 24) can be used.

[0050] The generator 21 is to be trained to map the latent space 20 to the ND space 22 of the system. The projection encoder 26 is to be trained to map the ND space 22 to the latent space 20. The loop formed between the generator 21 and the projection encoder 26 allows the training of one to benefit from the training of the other. The same samples or training data can be used to train the transformation between the configuration and the performance in both directions.

[0051] The generator 21 and the discriminator 24 form a GAN. The generator 21 learns to output performance, and the discriminator 24 learns to distinguish the performance, or classify the performance as real or fake (e.g., impossible versus operable or real). InFigure 2 In the arrangement, the anatomical structure model 23 intervenes such that the discriminator 24 receives the constrained or constrained-information performance as input, where the constrained information comes from the correlation between the performance of the robotic configuration and the anatomical structure. In an alternative embodiment, the performance is directly input to the discriminator without or in addition to the information from the anatomical structure model 23.

[0052] In one embodiment, the machine training includes constraints based on modeling the performance with respect to the anatomical structure. In operation 14, these constraints are determined. The configured performance is modeled with respect to the anatomical structure using the anatomical structure model 23. The anatomical structure model 23 is a virtual environment computational model that has its own latent space vector to describe the model as low-dimensional space parameters. The anatomical structure of interest (such as the liver or heart) is used to determine the constraints on the performance. For example, in a robot used to perform transrectal ultrasound, the trained computational model of the prostate includes parameterized tissue properties, target locations, the urethra, etc.

[0053] The anatomical structure model 23 is used in simulations or comparisons with the performance. For example, the anatomical structure model 23 defines the workspace or spatial extent of one or more components or the entire robot. As another example, the anatomical structure model 23 includes tissue information for determining the force levels that may damage the tissue. The performance represented in the ND space 22 is used to determine success, failure, and / or risk associated with the use of the configuration in the anatomical structure of interest at least for a given shape, position, and / or motion.

[0054] The anatomical structure model 23 uses anatomical representations from atlases, imaging data, and / or another source. For a fully automated system, the anatomical structure model 23 can be obtained from imaging data such as magnetic resonance (MR) data. For example, consider a robot used to perform transrectal ultrasound scans of the prostate. The workspace is derived by considering the MR segmentation of the prostate images of a population and determining the percentile of the reachable volume that the configured robot must satisfy. A statistical shape model is formed as the anatomical structure model from the imaging data. Similarly, the force application ability can be obtained by using a biomechanical model on these segmentations and determining the forces at each position in the workspace that will result in a fixed amount of strain in the prostate. A projection model is trained to map these ND spaces 22, either constrained or unconstrained by the anatomical structure model 23, to the latent space vector 20.

[0055] During training, the anatomical model 23 is used. When the generator 21 generates the performance for a given configuration, the constraints are determined for the performance of that configuration. This can be used in training as an input to the encoder 26 without user editing 25 to learn to generate configurations that are less likely to violate the constraints from the anatomical model 23, or can be used in testing to constrain the space of the generated solutions to those desired by the designer. Modeling the interaction of the robot with the anatomy based on the performance characteristics of the robot provides an adjusted value of the performance (i.e., a value that does not violate the constraints), which can be used as an input to the projection encoder. This constraint and / or the adjusted performance that satisfies the constraint is used as an input to the encoder 26 together with other performance information.

[0056] This machine training can include training the projection encoder 26 to project the performance and / or the adjusted performance into a configuration. The values of the ND space 22 output by the generator 21 and / or the anatomical model 23 are projected back into the latent space vector 20 to create another configuration. The projection encoder 26 is trained such that in an application, the user can then use channels for sensitivity or other performance to scale or adjust the amount across the workspace of the robot. These edits are then re-projected into the latent space 20 to produce an updated model of the robot, and the process can be repeated. The projection model is trained to produce the latent space vector 20 for a given input shape.

[0057] For training, the learnable parameters of the neural networks of the generator 21, discriminator 24, and / or projection encoder 26 are established. An optimization such as Adam is performed to learn the values of the learnable parameters. The learnable parameters are the weights, connections, convolutional kernels, and / or another characteristic of the neural network. Feedback for this optimization is provided by comparing the output of a given configuration using the neural network with the ground truth. For the generator 21, the output performance is compared with the known performance. For the projection encoder 26, the configuration output is compared with one or more known configurations for the input performance. Any difference function such as L1 or L2 can be used.

[0058] For the discriminator 24, any other classification as real or fake, physically reasonable or unreasonable, or distinguishing good or bad outputs from the generator 21 and / or the interaction with the anatomical model 23 is compared with the ground truth - real or fake. In the case of using a GAN, the output of the discriminator 24 can be used in the optimization of the generator 21 such that the generator 21 better learns to generate good outputs. The generator 21 and discriminator 24 are trained in an interleaved manner to improve the accuracy of both. The projection encoder 26 can benefit from this training of the GAN.

[0059] At Figure 1In operation 16, the learned or trained generator 21 and / or projection encoder 26 are stored. The generator 21 and encoder 26 are matrices or architectures having learned values for learnable parameters (e.g., convolutional kernels) and set values for other parameters. The machine learning network is stored for use in designing a robot according to modular components required for a given task. The discriminator 24 is not stored since the discriminator 24 is not used in the application. The discriminator 24 can be stored.

[0060] The learned network is stored in a memory together with the training data, or stored in other memories. For example, copies of the learned network are distributed to different computers or distributed across different computers for designing robots in different local environments. As another example, the copies are stored in the memory of one or more servers for online or cloud-based robot configuration.

[0061] Figure 6 A flowchart of an embodiment of a method for generating a medical robot according to configurable or reconfigurable modules is shown. The generative model 21 is trained such that an iterative method provides a resulting forward model to simulate an environment given a latent space vector 20 followed by the design of a projection model 26 that maps an ND space 22 of the design to the latent space vector 20, which is employed to refine the robot configuration. The generator 21 and / or projection encoder 26 are used in a single pass or used iteratively by repetition to test, design, and assemble a task-specific robot from standard or available robot components or modules. The memory stores the projection encoder 26 and / or generator 21 previously trained for the application.

[0062] The method is performed by a computer, such as Figure 7 system 70. Other computers, such as servers or workstations, can be used. For assembly, an automated or manual system follows the configuration to build a robot from component parts.

[0063] The operations are performed in the order shown (top to bottom or numerically) or in other orders. For example, the process can start with configuration, so generation at 64 is performed before any projection of capabilities at 62. As another example, user input at 67 is performed before operation 65.

[0064] Additional, different, or fewer actions may be used. For example, user inputs that do not perform action 67 and / or inputs of desired capabilities, such as where the automation system randomly or systematically designs different robots, or where the search system generates desired capabilities for a given task. As another example, the modeling of action 65 is not provided. In yet another example, one of actions 62 and 64 is not used, such as where the projection encoder 26 trained with GAN is used for configuration or design without inverse modeling to determine capabilities.

[0065] In action 60, the user inputs the capabilities of the medical robot. Using a user input device, such as a user interface with a keyboard, mouse, trackball, touchpad, touchscreen, and / or display device, the user inputs the desired capabilities of the medical robot.

[0066] Given the task for which the robot is to be designed and used, the desired capabilities are determined. Alternatively, the task is input and the processor extracts the capabilities. Any capabilities may be used, such as any information in the ND space 22. For example, the user inputs force, motion, compliance, workspace, payload, or joint requirements, constraints, goals, and / or tolerances. Logical, software, human-machine interface, and / or packaging capabilities may be input. The modeling of action 65 may be used to determine some, none, or all of the desired capabilities. By inputting capabilities rather than configurations, the user can avoid having to try to set various configuration variables (e.g., select which modules to use, how many of each module, interconnections, and / or adjustable parameters of each selected module).

[0067] In action 62, the processor uses or applies the machine learning encoder 26 to project the capabilities into the latent space vector 20, which defines the configuration of configurable modules (e.g., mechanical, software, human-machine interface, AI, and / or packaging configurable modules). Based on previous training, the encoder 26 outputs one or more configurations based on the input capabilities. The output is any latent space representation of the robot, such as a latent space vector, the shape of the assembled robot, or values of the types of configurable modules, the connections between configurable modules, and the parameters of the adjustable aspects of the configurable modules. Based on inputs of workspace, joint space, force space, compliance space, and / or kinematic space, the encoder 26 projects the user-defined capabilities into the configuration and parameter space of the configurable modules. Based on this projection, the configuration of the configurable modules is determined for the user-defined workspace, joint space, force space, compliance space, and / or kinematic space. Any ND space definition of performance can be transformed into a latent space definition of the configuration, such as the shape of the medical robot.

[0068] In operation 64, the processor uses or applies the machine learning generator 21 to generate an estimate of the capabilities. A latent space vector 20 such as output from the projection is input to the machine learning generator 21. The configured values are input. One or more settings of the configuration can be changed.

[0069] Based on this input, capabilities for the projected configuration are determined. These capabilities may or may not match the capabilities input to the projection encoder 26. The generator 21 generates estimates for the ND space 22, such as estimates for a workspace, a joint space, a force space, a compliance space, and / or a kinematic space. The output of the generator 21 in response to the input of the latent vector space 20 for this configuration is the capabilities of the ND space 22. In one embodiment, one or more of the capabilities are provided as three-dimensional vectors, such as forces in three dimensions. One or more output channels of the generator 21 can be three-dimensional outputs for modeling in three dimensions, such as for determining interactions with the anatomical model 23.

[0070] In one example embodiment, the machine learning network can provide optimal parameters to match the desired sensitivity in the robot pose. The pose estimation sensitivity is optimized by the GAN through appropriate reconfiguration of the module.

[0071] In operation 65, the processor models the operation of the medical robot of this configuration relative to the anatomical model 23. The capabilities with or without shape or other configuration information estimated by the generator 21 and / or input by the user are modeled relative to the anatomy. A computational model of the anatomy of interest (such as a computational model generated from imaging data and / or an atlas) is used to simulate the interaction of the robot with the tissue. The modeling can be of the robot in a given state, such as a shape having a given compliance, force, velocity, and / or other capabilities at a given position relative to the anatomy. The modeling can be of the robot for transitions between states (e.g., the dynamics or motion of the robot).

[0072] The modeling provides a design of the medical robot to interact with a complex three-dimensional medical environment. The modeling of the environment can be a machine learning model of the tissue, such as a biomechanical or computational model fit to the imaging data. Other modeling can be used, such as a three-dimensional mesh having elastic and / or other tissue properties assigned to the mesh. Medical imaging data can be used to fit or create a parameterized computational model of the anatomy. The anatomical model 23 is used with the configuration and performance to simulate the interaction.

[0073] In operation 66, the modeling is used to determine one or more constraints. The modeling indicates whether the capabilities exceed any limits. For example, compliance may result in poor positioning relative to the tissue. As another example, force may result in puncture, tearing, or undesired damage to the tissue. The anatomical model 23 is used to constrain the operation of the medical robot. The constraint can exist in any of the ND spatial variables, such as workspace, joint space, force space, compliance space, and / or kinematic space. The constraint can be a fixed limit or can be a goal.

[0074] Without violating the constraints, a feasible configuration is found. The process can proceed to operation 68 to complete the design of at least one possible configuration of the medical robot for the desired task.

[0075] In another embodiment, the constraint results in a change in one or more of the capabilities. For example, the force is reduced to avoid violating the constraint. The value of the capability is changed to be within the constraint. The changed value can then be used in the repetition of the projection of operation 62 to determine another configuration.

[0076] In operation 67, changes are made manually. A user input for a change in one or more of the estimated capabilities is received. The user edits one or more of the estimates. The edit can be to satisfy a constraint or for other reasons. For example, the user views the simulation or results from the modeling of operation 65 and decides to change the estimates.

[0077] Then, the estimates with any changes are projected, repeating operations 62 - 66 for that change. This iterative approach allows the refinement of the medical robot design based on using the input relative to performance rather than guessing changes in the configuration space.

[0078] User editing and subsequent projection, capability generation, and modeling from the edited capabilities allow for a balance of considerations of "similarity" (favoring consistency with user edits) and "feasibility" (favoring a feasible output). The balance can be aided by restricting the user input to one at a time or to only one of the estimated capabilities being changed for a given repetition or iteration in the design. For example, the user can only edit the workspace, force applicability, or payload capacity. Alternatively, the user can edit multiple estimates at once.

[0079] In another embodiment, the user is restricted to only editing capabilities. Direct editing or changing in the configuration is avoided. The user edits the capabilities to provide any changes in the configuration using the projection of operation 62. For example, instead of editing the shape, the user can only modify a scalar intensity proportional to a scalar (such as payload capacity) represented in a particular region of the workspace.

[0080] In one embodiment, the editing is an editing of the direction of the ability estimate. The user can edit the main directionality of the tensor channels, such as editing the compliance in a given direction. Consider a robot performing an ultrasound scan. The robot preferably has greater compliance along the ultrasound direction to avoid injury, while being relatively rigid in other directions. The user adjusts the compliance in the direction of the ultrasound scan line (e.g., perpendicular to the face of the ultrasound transducer).

[0081] Other edits can be made, such as allowing the user to edit one or more settings of the configuration (i.e., editing in the latent space). As an example, the user can switch the category (i.e., sub-module) label of a sub-module. The user selects different modules for a given part of the medical robot. As another example, the user can select how the robotic mechanism interacts with the target anatomical structure in the virtual space.

[0082] User editing can be provided based on the ability estimate, thus skipping the modeling in action 65. Alternatively, the modeling is used to inform the editing, such as providing constraint information to guide the user editing.

[0083] In another embodiment, the repetition or interaction is guided by another machine learning model. Deep reinforcement learning learns a policy for decision-making through a process of repetition and / or interaction. The policy is applied at each time or point in the process to select or set any variables, such as what performance characteristics to change and / or how much to change. The policy trained using deep reinforcement learning automates the repetition and / or interaction. The policy can be used to refine the design.

[0084] In action 68, the configuration of the configurable module is defined. The result of the modeling in action 65 may show a satisfactory design. The design may meet the constraints. The result is a configuration that can be built. Other satisfactory designs can be tried, so as to provide a choice of different options for the medical robot based on standardized or available modular components. The defined configuration is the final design to be manufactured.

[0085] In action 69, the robot is assembled. The robot, manufacturer, and / or designer can assemble the robot. The modular components to be used are collected, such as selecting the required quantity of each component included in the defined configuration. The modular components are configured as provided in the latent space vector 20. Then, the configured components are interconnected as provided in the latent space vector 20. Mechanical, software, circuit or electronic, human-machine interface, packaging, and / or other modules are assembled. The resulting medical robot designed for a specific task and corresponding anatomical structure can be used to assist in the diagnosis, treatment, and / or surgery of one or more patients.

[0086] Figure 7Illustrated is a system 70 for generating a medical robotic system according to reconfigurable modules. System 70 implements Figure 1 , Figure 2 , Figure 6 's method, or another method. In one embodiment, system 70 is for applications of a machine learning generation network 75 and / or a projection encoder network 76. Given input performance or task information, system 70 uses the encoder network 76 and / or the generation network 75 to design a medical robotic system. Although system 70 is described below in the context of applications of previously learned networks 75, 76, system 70 can be used to machine-train networks 75, 76 using many samples of robotic system designs for different capabilities and / or tasks.

[0087] System 70 includes an artificial intelligence processor 74, a user input 73, a memory 77, a display 78, and a medical scanner 72. The artificial intelligence processor 74, the memory 77, and the display 78 are shown as separate from the medical scanner 72, such as being part of a workstation, computer, or server. In an alternative embodiment, the artificial intelligence processor 74, the memory 77, and / or the display 78 are part of the medical scanner 72. In still other embodiments, system 70 does not include the medical scanner 72. Additional, different, or fewer components may be used.

[0088] The medical scanner 72 is a CT, MR, ultrasound, camera, or other scanner for scanning a patient. Scanner 72 may provide imaging data representative of one or more patients. The imaging data may be used for fitting and / or for creating biomechanical or computational models. The imaging data provides information for an anatomical model, thus allowing the robotic system to be modeled relative to the medical environment and tissue. The artificial intelligence processor 74 or other processor creates and / or uses the model to determine constraints on the robotic configuration and / or to refine the robotic configuration.

[0089] The memory 77 is a buffer, cache, RAM, removable media, hard drive, magnetic, optical, database, or other memory now known or later developed. The memory 77 is a single device, or a group of two or more devices. The memory 77 is shown as associated with or part of the artificial intelligence processor 74, but may be external to or remote from other components of system 70. For example, the memory 77 is a database storing many samples of robotic systems for use in training.

[0090] The memory 77 stores scan or image data, configurations, latent space vectors, ND space information, the machine learning generation network 75, the encoder network 76, and / or information for use in image processing to create a robotic system. For training, training data (i.e., input feature vectors and ground truth) is stored in the memory 77.

[0091] The memory 77 is additionally or alternatively a non-transitory computer-readable storage medium having processing instructions. The memory 77 stores data representing instructions executable by the programmed artificial intelligence processor 74. Instructions for implementing the processes, methods, and / or techniques discussed herein are provided on a computer-readable storage medium or memory such as a cache, buffer, RAM, removable media, hard drive, or other computer-readable storage medium. The machine learning generation network or image-to-image network 45 may be stored as part of the instructions for segmentation. Computer-readable storage media include various types of volatile and non-volatile storage media. The functions, acts, or tasks illustrated in the figures or described herein are performed in response to one or more sets of instructions stored in or on a computer-readable storage medium. These functions, acts, or tasks are independent of the particular type of instruction set, storage medium, processor, or processing strategy and may be performed by software, hardware, integrated circuits, firmware, microcode, etc., operating alone or in combination. Similarly, processing strategies may include multiprocessing, multitasking, parallel processing, etc.

[0092] In one embodiment, the instructions are stored on a removable media device for reading by a local or remote system. In other embodiments, the instructions are stored in a remote location for transmission over a computer network or over a telephone line. In still other embodiments, the instructions are stored in a given computer, CPU, GPU, or system.

[0093] The artificial intelligence processor 74 is a general-purpose processor, digital signal processor, three-dimensional data processor, graphics processing unit, application-specific integrated circuit, field-programmable gate array, digital circuit, analog circuit, massively parallel processor, a combination thereof, or other device now known or later developed for applying the machine learning networks 75, 76 and / or for performing modeling as part of a robotic design from modular components. The artificial intelligence processor 74 is a single device, multiple devices, or a network. For more than one device, parallel or sequential processing partitioning may be used. Different devices making up the artificial intelligence processor 74 may perform different functions. The artificial intelligence processor 74 is a hardware device configured by or operating according to stored instructions, designs (e.g., application-specific integrated circuits), firmware, or hardware to perform the various actions described herein.

[0094] The machine learning generator network 75 is trained to estimate the ability of the operations of the configuration of the robot. The machine learning encoder network 76 is trained to determine the configuration of the robot based on the ability. In an application, the artificial intelligence processor 74 uses the encoder network 76 with or without the generator network 75 to determine the configuration or the performance related to the task in a given task. The modeling for the interaction with the anatomical structure and / or the user input on the user input device can be used to change the task or the performance information during the determination of the configuration.

[0095] Based on past training, the machine learning networks 75, 76 are configured to output information. This training determines what output to provide in the case of a given previously unseen input. This training results in different operations of the networks 75, 76.

[0096] The user input 73 is a mouse, a trackball, a touchpad, a touch screen, a keyboard, a keypad, and / or other devices for receiving input from the user in the interaction with the computer or the artificial intelligence processor 74. The user input 73 forms a user interface with the display 78. The user input 73 is configured through the operating system to receive the input of the task or other performance information of the robot designed according to the reconfigurable components.

[0097] The display 78 is a CRT, an LCD, a plasma, a projector, a printer, or other output devices for showing the simulation of the interaction of the robot with the anatomical model, the robot configuration, and / or the robot performance. The display 78 is configured to display images through the image plane memory that stores the created images.

[0098] Although the present invention has been described above with reference to various embodiments, it should be understood that many changes and modifications can be made without departing from the scope of the present invention. Therefore, it is intended that the foregoing detailed description be considered illustrative rather than restrictive, and it is to be understood that it is the following claims, including all equivalents, that are intended to define the spirit and scope of the present invention.

Claims

1. A method for generating a medical robot according to configurable modules, the method comprising: Inputting (60) a first capability of the medical robot; Projecting (62) the first capability by a machine learning encoder into a latent space vector that defines (68) the configuration of the configurable module; Generating (64) an estimate of a second capability by a machine learning generator based on the input of the latent space vector to the machine learning generator; Modeling (65) the operation of the configured medical robot based on the second capability with respect to an anatomical structure model, wherein the modeling (65) is used to determine whether one or more constraints on the operation of the medical robot are satisfied; and Setting (68) the configuration of the configurable module based on the result of the modeling (65) and the latent space vector.

2. The method according to claim 1, wherein the input (60) comprises: Inputting (60) force, motion, compliance, workspace, load, and / or joint position.

3. The method according to claim 1, wherein the projection (62) comprises: Projecting (62) from the first capability into the latent space vector, the latent space vector including values of parameters for the type of the configurable module, connections between the configurable modules, and adjustable aspects of the configurable module.

4. The method according to claim 1, wherein generating (64) comprises: Generating (64) by a machine learning generator that has been trained as a generative adversarial network.

5. The method according to claim 1, wherein generating (64) comprises: Generating (64) by a machine learning generator including a neural network.

6. The method according to claim 1, wherein generating (64) comprises: Generating (64) an estimate of the second capability as a three-dimensional capability.

7. The method according to claim 1, further comprising: Receiving (67) a user input for changing one of the estimates of the second capabilities, repeating the projection (62) using the estimate of the second capabilities including the changed estimate, repeating the generating (64) according to another latent space vector projected by the repetition of the projection (62), and repeating the modeling (65) based on a third capability obtained from the repetition of the generating (64).

8. The method according to claim 7, wherein receiving (67) the user input comprises: Restricting the user input to changing only one of the estimates of the second capabilities.

9. The method according to claim 7, wherein receiving (67) a user input comprises: Receiving (67) a user input for changing one estimate without changing the configuration.

10. The method according to claim 7, wherein receiving (67) the user input comprises: Receiving (67) a change as a change in the direction of the estimate of the second capability.

11. The method according to claim 1, wherein the modeling (65) comprises: Modeling (65) using the anatomical structure model, the anatomical structure model including a trained computational model from image data.

12. The method according to claim 1, wherein the projection (62) comprises: Performing the projection (62) in the case where the configuration includes the shape of the medical robot.

13. A method for machine training (12) to configure a robot according to component modules, the method comprising: Providing (10) training data, the training data including various configurations of the robot from component models and the performance of the robot; Machine training (12) a generator of a generative adversarial network using the training data to estimate performance based on the input of the configuration, the generative adversarial network further including a discriminator configured to map the performance to a classification of real or fake, wherein the machine training (12) includes constraints based on modeling (65) the performance with respect to an anatomical structure; Storing (16) the machine-trained generator of the generative adversarial network; The machine training (12) includes: by constraints of modeling (65) the configuration relative to the anatomical structure, the modeling (65) provides (10) an adjusted value of the performance, and machine training (12) the encoder to project the performance and / or adjusted performance onto the configuration.

14. The method according to claim 13, wherein modeling the configuration (65) relative to the anatomical structure comprises: Modeling (65) is performed using an anatomical structure model from medical imaging.

15. A method for generating according to a configurable module of a medical robot, the method comprising: Projecting (62) from a user-defined workspace, joint space, force space, compliance space, and / or kinematic space to the configuration and parameter space of the configurable module, the projection (62) being performed by a machine learning encoder that has been trained based on a generative adversarial network; Determining (68) the configuration of the configurable module for the user-defined workspace, joint space, force space, compliance space, and / or kinematic space according to the projection (62); Estimating values in the workspace, joint space, force space, compliance space, and / or kinematic space by the generative adversarial network from the configuration and parameter space of the configurable module; Wherein the method further includes: modeling (65) the interaction with the anatomical structure based on the estimated values in the workspace, joint space, force space, compliance space, and / or kinematic space, thereby constraining (66) the user-defined workspace, joint space, force space, compliance space, and / or kinematic space.

16. The method according to claim 15, further comprising: Generating (64) an estimate of the workspace, joint space, force space, compliance space, and / or kinematic space by a machine learning generator of the generative adversarial network from the values of the configuration and parameter space of the configurable module; receiving (67) user editing of the estimate; and Repeating the projection from the edited estimate.

Citation Information

Patent Citations

  • Solution method of joint space parameters of artificial limb in multiple degrees of freedom

    CN101953727A

  • Method and system for guiding user positioning of robot

    CN108778179A

  • Generative neural network systems for generating instruction sequences to control an agent performing a task

    WO2019155052A1