Ultrasonic probe compliance control method and system based on inverse reinforcement learning
Through inverse reinforcement learning and dynamic inverse solution algorithm, the contact force and attitude control of the ultrasonic probe are optimized, and the problem of posture instability of the ultrasonic probe during skin contact is solved, achieving high-quality ultrasonic imaging.
Patent Information
- Application Number
- CN202510353093.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to achieve precise control of the stable posture and appropriate contact force of ultrasonic probes during skin contact, resulting in unstable imaging quality.
Using an inverse reinforcement learning method, the strategic network is optimized by designing reasonable reward functions and expert demonstration data, establishing a mapping relationship between contact force and probe attitude, and combining dynamic inverse solution algorithm to control the movement of the robotic arm to achieve the smooth control of the ultrasonic probe.
It realizes precise control of ultrasonic probe attitude adjustment and contact force, improving the stability and reliability of ultrasonic image quality.
Smart Images

Figure CN120267330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robotic breast ultrasound, and more particularly, to a compliant control method and system for an ultrasound probe based on inverse reinforcement learning. Background Art
[0002] Ultrasound scanning robots have become one of the core topics in the field of robotics research due to their advantages of automatically performing repetitive operations and effectively reducing the burden on operators.
[0003] In autonomous ultrasound scanning tasks, the contact force and probe attitude directly determine the imaging quality and the reliability of the diagnostic results. To obtain clear and high-quality ultrasound images, the probe usually needs to maintain a stable attitude perpendicular to the skin surface, as Figure 1 shown, while applying an appropriate contact force. However, during the contact process between the probe and the skin, the skin exhibits complex non-linear deformation and stiffness changes, which significantly increase the difficulty of adjusting the probe attitude and controlling the contact force. Different deformation depths and probe attitudes will result in significant differences in the contact force. This complexity poses a severe challenge to achieving stable robotic ultrasound scanning control. Summary of the Invention
[0004] The purpose of the present invention is to provide a compliant control method and system for an ultrasound probe based on inverse reinforcement learning. By designing a reasonable reward function and combining reinforcement learning methods, using the ultrasound image quality as the reward criterion, initially establishing the mapping relationship between the contact force and the probe attitude, and performing secondary optimization on the policy network through the inverse reinforcement learning method and expert demonstrations, precise control of the ultrasound probe attitude adjustment and contact force is achieved, aiming to solve the problems in the prior art.
[0005] The present invention is implemented as follows. The compliant control method for an ultrasound probe based on inverse reinforcement learning is applied to a breast ultrasound controller and specifically includes the following steps:
[0006] S101: Build and configure a neural network model for reinforcement learning on a deep learning platform, configure the reward function, determine whether to give a reward by judging whether ultrasound imaging is detected, and update the network parameters online to complete the first training of the network;
[0007] S102: Collect the scanning trajectories demonstrated by human experts under different deformation depths and skin stiffness conditions, as well as the contact force and attitude adjustment data of the ultrasound probe;
[0008] S103: Perform secondary optimization of the policy network based on inverse reinforcement learning and expert demonstration data, configure the structure of the generative adversarial inverse reinforcement learning network, and respectively configure the discriminator and the policy network. Learn the implicit reward function through the optimal policy to complete the second training of the policy network;
[0009] S104: Implement the contact force and pose control of the autonomous ultrasonic scanning robot using the trained policy network, and control the movement of the robotic arm through the inverse dynamics algorithm to transform the ultrasonic probe to the desired pose.
[0010] Further, in S101, configure the reward function, and determine whether to give a reward by judging whether ultrasonic imaging is detected, including:
[0011] When complete ultrasonic imaging is detected, give a fixed positive reward R s ;
[0012] If ultrasonic imaging is not detected or the imaging is incomplete, give a fixed negative reward R p 。
[0013] Further, and update the network parameters online to complete one training of the network, including:
[0014] Generate quadruples in real time from the environmental interaction, that is, data of state, action, reward, and next state, and perform normalization operations;
[0015] According to the reinforcement learning algorithm, select an appropriate loss function, and the loss function is a policy network loss function or a value function loss function;
[0016] Continuously collect new environmental interaction data and add it to the data set, and iteratively update the neural network model online.
[0017] Further, in S102, collect the scanning trajectories demonstrated by human experts under different deformation depths and skin stiffness conditions, including:
[0018] Set skin models with different stiffness and deformation characteristics, and adjust the experimental environment;
[0019] When the human expert manually operates the ultrasonic probe to perform the scanning task, record the pose of the ultrasonic probe and the contact force information of the 6-degree-of-freedom force sensor in real time;
[0020] Identify the recorded pose of the ultrasonic probe and the contact force information of the 6-degree-of-freedom force sensor, obtain the outlier information, and complete the cleaning of the outlier information.
[0021] Further, in S103, perform secondary optimization of the policy network based on inverse reinforcement learning and expert demonstration data, including:
[0022] Configure the structure of the generative adversarial inverse reinforcement learning network, and configure the discriminator and the policy network respectively;
[0023] Learn an implicit reward function from expert demonstration data. Through adversarial training on the scanning trajectories of human expert demonstrations and the trajectories generated by the policy, improve the fitting ability of the reward function to expert behavior, obtain the classification probability that maximizes the positive samples, and at the same time minimize the classification probability of negative samples, making the reward function closer to the reward distribution of expert behavior.
[0024] Furthermore, the discriminator is used to distinguish whether a sample belongs to expert demonstration. Its output L represents the probability that the sample belongs to expert demonstration, and the loss function of the discriminator D(s, a) is defined as:
[0025]
[0026] Furthermore, in S103, the optimization of the policy network is completed, including:
[0027] Load the policy network π θ (s) obtained from the initial reinforcement learning training of the neural network model as the starting point of optimization, and combine it with the implicit reward function obtained through inverse reinforcement learning Configure the optimization objective function J(θ), which represents the cumulative reward of the policy in the task
[0028] The optimization goal is to adjust the policy network parameters θ to maximize the cumulative reward of the policy;
[0029] Complete the real-time interaction between the policy network and the environment to generate trajectory data (s, a, r, s′), where the reward r is dynamically provided by the reward function. According to the sampled trajectory data, use the optimization algorithm to update the policy network parameters θ through gradient descent, adjust the behavior of the policy, and through multiple iterations of trajectory collection and parameter update, the policy network gradually learns to adapt to the complex environment.
[0030] Furthermore, in S104, the contact force and pose control of the autonomous ultrasonic scanning robot are realized for the trained policy network, including:
[0031] Load the trained policy network π θ (s). The policy network π θ (s) has learned the optimized control strategy for contact force and probe pose. The input of the policy network is the current 6-degree-of-freedom contact force information, and the output is the action space, including the target contact force F target in the contact direction, the target torques τ roll 、τ pitch in the rolling direction and pitching direction, and the corresponding linear velocity and angular velocity
[0032] The state data obtained in real time is input into the trained policy network, and the next action space is output, and this action space is sent to the breast ultrasound controller;
[0033] After the breast ultrasound controller receives the action space information generated by the policy network, it controls the movement of the robotic arm through the inverse dynamics algorithm to transform the ultrasound probe to the desired pose.
[0034] Compared with the prior art, the ultrasound probe compliance control method and system based on inverse reinforcement learning provided by the present invention have the following beneficial effects:
[0035] 1. First, use reinforcement learning to train the policy network once, and then introduce expert demonstration data and inverse reinforcement learning to further optimize the policy network twice. The specific implementation of each link includes: performing initial policy training, collecting expert demonstration data, secondary optimization, and adjusting the pose based on the policy network. By designing a reasonable reward function and combining the reinforcement learning method, using the ultrasound image quality as the reward standard, initially establishing the mapping relationship between the contact force and the probe pose, and further optimizing the policy network through the inverse reinforcement learning method and expert demonstration, realizing the precise control of the ultrasound probe pose adjustment and contact force.
[0036] 2. The core is to construct a robust coupling relationship between the ultrasound scanning robot and the skin, which can simultaneously realize contact force adjustment and precise control of the probe pose. By sensing the 6-degree-of-freedom force information, the pose and force control instructions for the next moment are deduced. This method first uses reinforcement learning to initially optimize the policy network, and then introduces expert demonstration data and inverse reinforcement learning to further optimize the policy network twice. Finally, the policy network receives the data of the 6-degree-of-freedom force sensor to optimize the pose adjustment and contact force control of the ultrasound probe. In the second round of optimization, an inverse reinforcement learning framework is introduced, and the implicit reward function is learned through expert demonstration data to further improve the adaptability and robustness of the policy.
[0037] An ultrasound probe compliance control system based on inverse reinforcement learning is used to execute the above-mentioned ultrasound probe compliance control method. The control system includes:
[0038] A primary training module for building and configuring a neural network model of reinforcement learning on a deep learning platform, and configuring a reward function to complete the primary training of the network;
[0039] A data acquisition module for collecting the scanning trajectories demonstrated by human experts, as well as the contact force and pose adjustment data of the ultrasound probe;
[0040] A secondary optimization module for configuring a generative adversarial inverse reinforcement learning network structure, and respectively configuring a discriminator and a policy network, obtaining an implicit reward function through inverse reinforcement learning, and completing the secondary optimization of the network;
[0041] The inverse solution execution module is used to control the contact force and pose of the autonomous ultrasonic scanning robot for the trained policy network, and control the movement of the robotic arm through the dynamic inverse solution algorithm to transform the ultrasonic probe to the desired pose. Description of the Drawings
[0042] Figure 1 It is a schematic diagram of the detection pose of the ultrasonic probe in the prior art;
[0043] Figure 2 It is a schematic flowchart of the ultrasonic probe compliance control method based on inverse reinforcement learning proposed by the present invention;
[0044] Figure 3 It is a logical operation diagram of the ultrasonic probe compliance control method based on inverse reinforcement learning proposed by the present invention;
[0045] Figure 4 It is a schematic structural diagram of the ultrasonic probe compliance control system based on inverse reinforcement learning proposed by the present invention. Detailed Embodiment
[0046] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0047] The implementation of the present invention will be described in detail below with reference to specific embodiments.
[0048] In the drawings of this embodiment, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as limitations of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0049] Refer to Figure 2-3 As shown, the ultrasonic probe compliance control method based on inverse reinforcement learning is applied to a breast ultrasound controller, and specifically includes the following steps:
[0050] S101: Build and configure a neural network model for reinforcement learning on a deep learning platform, configure a reward function, determine whether to give a reward by judging whether ultrasonic imaging is detected, and update the network parameters online to complete one training of the network;
[0051] Among them, configure the reward function, and decide whether to give a reward by judging whether ultrasonic imaging is detected, including:
[0052] When complete ultrasonic imaging is detected, give a fixed positive reward R s ;
[0053] If ultrasonic imaging is not detected or the imaging is incomplete, give a fixed negative reward R p ;
[0054] Specifically, on a deep learning platform such as PyTorch, TensorFlow, etc., build and configure a neural network model for reinforcement learning, such as classic network structures like PPO or DDPG;
[0055] S102: Collect the scanning trajectories demonstrated by human experts under different deformation depths and skin stiffness conditions, as well as the contact force and pose adjustment data of the ultrasonic probe;
[0056] Among them, collecting the scanning trajectories demonstrated by human experts under different deformation depths and skin stiffness conditions includes:
[0057] Set up skin models with different stiffness and deformation characteristics, and adjust the experimental environment;
[0058] When a human expert manually operates the ultrasonic probe to perform a scanning task, record the pose of the ultrasonic probe and the contact force information of the 6-degree-of-freedom force sensor in real time;
[0059] Identify the recorded pose of the ultrasonic probe and the contact force information of the 6-degree-of-freedom force sensor, obtain outlier information, and complete the cleaning of outlier information;
[0060] Outliers are abnormal data points that do not conform to physical laws or statistical characteristics, usually caused by factors such as sensor noise and environmental interference.
[0061] S103: Based on inverse reinforcement learning and expert demonstration data, perform secondary optimization of the policy network, configure a generative adversarial inverse reinforcement learning network structure, and configure a discriminator and a policy network respectively. Learn the implicit reward function through the optimal policy and complete the secondary training of the policy network;
[0062] Among them, based on inverse reinforcement learning and expert demonstration data, performing secondary optimization of the policy network includes:
[0063] Configure a generative adversarial inverse reinforcement learning network structure, and configure a discriminator and a policy network respectively;
[0064] Learn an implicit reward function from expert demonstration data. Through adversarial training of the scanning trajectory of human expert demonstrations and the trajectory generated by the policy, improve the fitting ability of the reward function to expert behavior, obtain the classification probability that maximizes the positive samples, and minimize the classification probability of negative samples at the same time, making the reward function closer to the reward distribution of expert behavior;
[0065] S104: Use the trained policy network to implement the contact force and pose control of the autonomous ultrasound scanning robot. Control the movement of the robotic arm through the inverse dynamics algorithm to transform the ultrasound probe to the desired pose;
[0066] Among them, using the trained policy network to implement the contact force and pose control of the autonomous ultrasound scanning robot includes:
[0067] Load the trained policy network. The policy network has learned the optimized control strategy for contact force and probe pose. The input of the policy network is the current 6-degree-of-freedom contact force information, and the output is the action space, including the target contact force F in the contact direction target and the target torques τ in the rolling direction and pitch direction roll 、τ pitch and the corresponding linear velocity and angular velocity
[0068] Input the real-time acquired state data into the trained policy network, output the next action space, and send the action space to the breast ultrasound controller;
[0069] After the breast ultrasound controller receives the action space information generated by the policy network, it controls the movement of the robotic arm through the inverse dynamics algorithm to transform the ultrasound probe to the desired pose. First, use reinforcement learning to preliminarily optimize the policy network, and then further perform secondary optimization on the policy network by introducing expert demonstration data and inverse reinforcement learning. The specific implementation of each link includes: performing initial policy training, collecting expert demonstration data, secondary optimization, and adjusting the pose based on the policy network. By designing a reasonable reward function and combining reinforcement learning methods, using the ultrasound image quality as the reward criterion, initially establish the mapping relationship between the contact force and the probe pose, and perform secondary optimization on the policy network through the inverse reinforcement learning method and expert demonstrations, realizing the precise control of the ultrasound probe pose adjustment and contact force.
[0070] In S101 of this embodiment, the online network parameters are updated, including:
[0071] Generate quadruples in real time from the environment interaction, that is, data of state, action, reward, and next state, and perform normalization operations;
[0072] According to the reinforcement learning algorithm, select an appropriate loss function. The loss function is the policy network loss function or the value function loss function;
[0073] Continuously collect new environmental interaction data and add it to the dataset, and iteratively update the neural network model online.
[0074] In S103 of this embodiment, the discriminator is used to distinguish whether a sample belongs to an expert demonstration. Its output L represents the probability that the sample belongs to an expert demonstration, and the loss function of the discriminator D(s, a) is defined as:
[0075]
[0076] In S103 of this embodiment, the secondary optimization of the policy network is completed, including:
[0077] Load the policy network π θ (s) obtained from the initial reinforcement learning training of the neural network model as the starting point of optimization, and combine the implicit reward function obtained through inverse reinforcement learning Configure the optimization objective function J(θ), which represents the cumulative reward of the policy in the task
[0078] The optimization objective is to adjust the policy network parameters θ to maximize the cumulative reward of the policy;
[0079] Complete the real-time interaction between the policy network and the environment to generate trajectory data (s, a, r, s′), where the reward r is dynamically provided by the reward function According to the sampled trajectory data, use the optimization algorithm to update the policy network parameters θ through gradient descent, adjust the behavior of the policy, and through multiple iterations of trajectory collection and parameter update, the policy network gradually learns to adapt to the complex environment.
[0080] The core of this technical solution lies in constructing a robust coupling relationship between the ultrasonic scanning robot and the skin, which can simultaneously achieve contact force adjustment and precise probe attitude control. Derive the pose and force control instructions for the next moment by sensing 6-degree-of-freedom force information. This method first uses reinforcement learning to preliminarily optimize the policy network, and then further performs secondary optimization on the policy network by introducing expert demonstration data and inverse reinforcement learning. Finally, it is realized that the policy network optimizes the pose adjustment and contact force control of the ultrasonic probe by receiving the data of the 6-degree-of-freedom force sensor. In the second round of optimization, an inverse reinforcement learning framework is introduced to learn the implicit reward function through expert demonstration data, further improving the adaptability and robustness of the policy.
[0081] Specifically, from Figure 3From the logical operation diagram, it can be seen that the 6-DOF force sensor generates an action space through the inner loop - force / pose adjustment module and sends it to the controller for execution. In the inner loop - force / pose adjustment module, the model is constructed, trained, and optimized. First, the policy network is optimized once based on the reinforcement learning method, and then through demonstrations with human experts and inverse reinforcement learning, the secondary optimization of the policy network is completed, realizing the contact force and pose control of the autonomous ultrasonic scanning robot. The robotic arm is controlled to move through the inverse dynamics algorithm to transform the ultrasonic probe to the desired pose, thus achieving the compliant control of the ultrasonic probe pose.
[0082] Referring to Figure 4 As shown, the compliant control system for the ultrasonic probe based on inverse reinforcement learning is used to execute the above-mentioned compliant control method for the ultrasonic probe. The control system includes: a primary training module for building and configuring the neural network model of reinforcement learning on the deep learning platform and configuring the reward function to complete the primary training of the network; a data acquisition module for collecting the scanning trajectories demonstrated by human experts, as well as the contact force and pose adjustment data of the ultrasonic probe; a secondary optimization module for configuring the structure of the generative adversarial inverse reinforcement learning network, respectively configuring the discriminator and the policy network, and obtaining the implicit reward function through inverse reinforcement learning to complete the secondary optimization of the network; an inverse solution execution module for achieving the contact force and pose control of the autonomous ultrasonic scanning robot for the trained policy network, controlling the robotic arm to move through the inverse dynamics algorithm to transform the ultrasonic probe to the desired pose. By first using reinforcement learning to preliminarily optimize the policy network, and then introducing expert demonstration data and inverse reinforcement learning to further optimize the policy network, the specific implementation of each link includes: performing initial policy training, collecting expert demonstration data, secondary optimization, and adjusting the pose based on the policy network. By designing a reasonable reward function and combining the reinforcement learning method, using the ultrasonic image quality as the reward criterion, the mapping relationship between the contact force and the probe pose is initially established, and the secondary optimization of the policy network is carried out through the inverse reinforcement learning method and expert demonstrations, realizing the precise control of the ultrasonic probe pose adjustment and contact force.
[0083] In this embodiment, the entire operation process can be controlled by a computer, and in each operation link, signal feedback can be carried out by setting sensors to achieve the sequential execution of steps. These are all common knowledge in current automatic control and will not be elaborated one by one in this embodiment.
[0084] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An ultrasonic probe compliant control method based on inverse reinforcement learning, characterized in that, Applied to a breast ultrasound controller, specifically including the following steps: S101: Build and configure a neural network model for reinforcement learning on a deep learning platform, configure a reward function, determine whether to give a reward by judging whether ultrasound imaging is detected, and update network parameters online to complete one training of the network; S102: Collect the scanning trajectories demonstrated by human experts under different deformation depths and skin stiffness conditions, as well as the contact force and pose adjustment data of the ultrasound probe; S103: Perform secondary optimization of the policy network based on inverse reinforcement learning and expert demonstration data, configure the structure of a generative adversarial inverse reinforcement learning network, and configure a discriminator and a policy network respectively. Learn an implicit reward function through optimal policy learning to complete the secondary training of the policy network; S104: Use the trained policy network to achieve the contact force and pose control of an autonomous ultrasound scanning robot, and control the movement of the robotic arm through a kinematic inverse solution algorithm to transform the ultrasound probe to the desired pose.
2. The method for compliant control of an ultrasonic probe based on inverse reinforcement learning according to claim 1, wherein In S101, configure the reward function, and determine whether to give a reward by judging whether ultrasound imaging is detected, including: When a complete ultrasound image is detected, a fixed positive reward R is given s ; If no ultrasonic imaging is detected or the imaging is incomplete, a fixed negative reward R is given p .
3. The method for compliant control of an ultrasonic probe based on inverse reinforcement learning according to claim 2, characterized in that, And update network parameters online to complete one training of the network, including: Generate a quadruple (s, a, r, s') in real time from the environmental interaction, that is, data of state, action, reward, and next state, and perform a normalization operation; According to the reinforcement learning algorithm, select an appropriate loss function, and the loss function is a policy network loss function or a value function loss function; Continuously collect new environmental interaction data and add it to the dataset, and iteratively update the neural network model online.
4. The method for compliant control of an ultrasonic probe based on inverse reinforcement learning according to claim 3, characterized in that, In S102, collect the scanning trajectories demonstrated by human experts under different deformation depths and skin stiffness conditions, including: Set skin models with different stiffnesses and deformation characteristics, and adjust the experimental environment; When a human expert manually operates the ultrasound probe to perform a scanning task, record the pose of the ultrasound probe and the contact force information of the 6-degree-of-freedom force sensor in real time; Identify the recorded pose of the ultrasound probe and the contact force information of the 6-degree-of-freedom force sensor, obtain outlier information, and complete the cleaning of outlier information.
5. The method for compliant control of an ultrasonic probe based on inverse reinforcement learning according to claim 4, wherein In S103, perform secondary optimization of the policy network based on inverse reinforcement learning and expert demonstration data, including: Configure the structure of a generative adversarial inverse reinforcement learning network, and configure a discriminator and a policy network respectively; Learn an implicit reward function from the expert demonstration data. Through adversarial training of the scanning trajectories demonstrated by human experts and the trajectories generated by the policy, improve the fitting ability of the reward function to expert behavior, obtain the classification probability of maximizing positive samples, and at the same time minimize the classification probability of negative samples, so that the reward function is closer to the reward distribution of expert behavior.
6. The method for compliant control of an ultrasonic probe based on inverse reinforcement learning according to claim 5, wherein The discriminator is used to distinguish whether a sample belongs to an expert demonstration. Its output L represents the probability that the sample belongs to an expert demonstration, and the loss function of the discriminator D(s, a) is defined as:
7. The method for compliant control of an ultrasonic probe based on inverse reinforcement learning according to claim 6, wherein In S103, complete the optimization of the policy network, including: Load the policy network π obtained from the initial reinforcement learning training of the neural network model θ (s), as the starting point for optimization, combined with the implicit reward function obtained through inverse reinforcement learning Configure the optimization objective function J(θ), which represents the cumulative reward of the policy in the task The optimization goal is to adjust the policy network parameters θ to maximize the cumulative reward of the policy; Complete the real-time interaction between the policy network and the environment to generate trajectory data (s, a, r, s'), where the reward r is provided dynamically by the reward function According to the sampled trajectory data, use the optimization algorithm to update the policy network parameters θ through gradient descent, adjust the behavior of the policy. Through multiple iterations of trajectory collection and parameter update, the policy network gradually learns to adapt to the complex environment.
8. The method for compliant control of an ultrasound probe based on inverse reinforcement learning according to claim 7, wherein In S104, achieve the contact force and pose control of an autonomous ultrasound scanning robot for the trained policy network, including: Load the trained policy network π θ (s), the policy network π θ (s) has learned the optimal control strategy for the contact force and the probe pose. The input of the policy network is the current 6-degree-of-freedom contact force information, and the output is the action space, including the target contact force F target in the contact direction, and the target torques τ roll 、τ pitch in the rolling and pitching directions, and the corresponding linear and angular velocities The status data obtained in real time is input into the trained policy network, and the next action space is output, and this action space is sent to the breast ultrasound controller; After receiving the action space information generated by the policy network, the breast ultrasound controller controls the movement of the robotic arm through the inverse kinematics algorithm to transform the ultrasound probe to the desired pose.
9. An ultrasonic probe compliant control system based on inverse reinforcement learning, characterized in that A control system for implementing the ultrasound probe compliance control method according to any one of claims 1-8, the control system comprising: A primary training module for building and configuring a neural network model for reinforcement learning on a deep learning platform, and configuring a reward function to complete the primary training of the network; A data acquisition module for collecting the scanning trajectories demonstrated by human experts, as well as the contact force and pose adjustment data of the ultrasound probe; A secondary optimization module for configuring a generative adversarial inverse reinforcement learning network structure, and respectively configuring a discriminator and a policy network, and obtaining an implicit reward function through inverse reinforcement learning to complete the secondary optimization of the network; An inverse solution execution module for implementing the contact force and pose control of the autonomous ultrasound scanning robot for the trained policy network, and controlling the movement of the robotic arm through the inverse kinematics algorithm to transform the ultrasound probe to the desired pose.