Kidney autonomous ultrasonic navigation method based on deep reinforcement learning

By training the DQN model through deep reinforcement learning, the ultrasound probe can autonomously locate the standard imaging plane in a virtual environment, solving the problems of time-consuming and labor-intensive technologies and poor imaging stability, and improving the efficiency and accuracy of renal ultrasound examinations.

CN120605102AInactive Publication Date: 2025-09-09FUDAN UNIV YIWU RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511093698.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing ultrasound examinations are highly dependent on operator experience in renal imaging, which is time-consuming and labor-intensive, and can easily lead to decreased imaging stability due to operator fatigue, making it difficult to quickly and accurately find the standard imaging plane.

Method used

Using deep reinforcement learning methods, a three-dimensional voxel point cloud virtual environment is constructed to train the DQN model. The virtual ultrasound probe is controlled to learn to autonomously locate the standard imaging plane in the virtual environment. The training results are then applied to the actual ultrasound probe to achieve autonomous navigation.

Benefits of technology

It significantly reduces doctors' manual operation time, improves the efficiency and stability of renal ultrasound examinations, ensures image quality, reduces repeated scans, and provides a stable foundation for accurate diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120605102A_ABST
    Figure CN120605102A_ABST
Patent Text Reader

Abstract

The invention relates to the field of medical assistance, in particular to a kidney autonomous ultrasonic navigation method based on deep reinforcement learning, which comprises the following steps: constructing a phantom three-dimensional virtual environment containing the kidney, constructing an intelligent agent for controlling a virtual ultrasonic probe in the phantom three-dimensional virtual environment, and performing ultrasonic navigation on the virtual ultrasonic probe. Configuring a kidney to a region of interest in the phantom three-dimensional virtual environment, and setting a target standard imaging plane; inputting the ultrasonic images of a plurality of recent time steps into a pre-trained UNet network, extracting feature vectors, and taking the feature vectors as input states of a DQN model; training the DQN model, and finally forming the DQN model capable of automatically positioning a standard imaging plane; and using the trained DQN model to control the movement of an actual ultrasonic probe. In the navigation process, the imaging quality is evaluated in real time, the probe pose is dynamically adjusted, the acquisition success rate of qualified images is improved, repeated scanning caused by insufficient image quality is effectively reduced, and a stable image basis is provided for accurate diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical assistance, and in particular to a kidney autonomous ultrasound navigation method based on deep reinforcement learning. Background Art

[0002] Ultrasound examination, as a commonly used medical imaging diagnostic technology in clinical practice, has significant advantages such as being non-invasive, real-time imaging, and low cost, and is widely used in the initial screening and diagnosis of kidney diseases. The core of renal ultrasound examination is to obtain standard imaging planes (such as longitudinal, transverse, and coronal sections) to clearly display anatomical structures such as the renal cortex, medulla, and renal sinus, as well as their adjacent relationships. The imaging quality directly affects the accuracy of diagnosis. However, current ultrasound examinations are highly dependent on the operator's experience. Doctors need to manually translate and rotate the ultrasound probe, repeatedly adjusting its position to find the standard plane. This process is time-consuming and labor-intensive, and can easily lead to a decrease in imaging stability due to operator fatigue. Summary of the Invention

[0003] In order to solve the problems mentioned in the above background technology, the present invention provides a kidney autonomous ultrasound navigation method based on deep reinforcement learning.

[0004] The present invention provides a deep reinforcement learning-based autonomous ultrasound navigation method for the kidney, which adopts the following technical solution: comprising the steps of: A 3D virtual environment containing the kidney is constructed using a 3D voxel point cloud. An intelligent agent is constructed in the 3D virtual environment to control a virtual ultrasound probe. The kidney is positioned within the 3D virtual environment to locate the region of interest, and a target standard imaging plane is set. The state of the agent is defined to consist of the posture of the virtual ultrasound probe and the ultrasound image features, and the virtual ultrasound probe is defined to be able to perform several discrete actions; Prepare a pre-trained UNet network and DQN model. Collect the current position of the virtual ultrasound probe and ultrasound image at each time step. Input the ultrasound images of the most recent several time steps into the pre-trained UNet network and extract the feature vectors, which serve as the input state of the DQN model. Train the DQN model, store historical states, actions, and reward data in the experience replay pool, and regularly synchronize the main network parameters with the target network for training, ultimately forming a DQN model that can autonomously locate the standard imaging plane; The trained DQN model is used to control the movement of the actual ultrasound probe.

[0005] Furthermore, a three-dimensional virtual environment of the kidney phantom is constructed using medical imaging data to construct a hierarchical voxel point cloud model of the anatomical structures of the cortex, medulla, and renal sinus. In the three-dimensional virtual environment, the position of the kidney is fixed at the origin of the virtual coordinate system and the surrounding area is filled with bionic tissue media. The acoustic properties of the bionic tissue medium match the properties of human soft tissue, simulating the propagation and attenuation characteristics of ultrasonic waves to generate dynamic ultrasonic echo signals.

[0006] Furthermore, the intelligent agent includes a perception unit, a decision unit and a memory unit. The perception unit is used to obtain the spatial posture information of the virtual ultrasound probe in real time, and at the same time collect ultrasound image data of the corresponding position of the virtual ultrasound probe; the decision unit integrates the DQN model, and the decision unit uses the posture and ultrasound image features output by the perception unit as input states, calculates the Q value of each discrete action to select the optimal action; the memory unit is used to store historical states, actions, and reward data in the experience replay pool to provide data support for the training optimization of the DQN model.

[0007] Furthermore, the region of interest is set as a cubic space containing the complete anatomical structure of the kidney, and the region of interest includes the renal cortex, medulla, and renal sinus tissue.

[0008] Furthermore, the target standard imaging planes are set based on typical standard sections used in clinical renal ultrasound examinations. The target standard imaging planes include three core planes: longitudinal, transverse, and coronal. Each plane definition is aligned with the virtual coordinate system. The longitudinal section is parallel to the z-axis and y-axis of the virtual coordinate system, with its normal direction being the x-axis of the virtual coordinate system. The longitudinal section fully displays the long-axis anatomical structure of the kidney, including the hilum, renal sinus, cortex, and medulla. Furthermore, the transverse section is perpendicular to the z-axis of the virtual coordinate system, parallel to the xy plane of the virtual coordinate system, and with the normal direction being the z-axis of the virtual coordinate system. The transverse section presents a short-axis cross-section of the kidney, showing the annular distribution of the high-echo area of ​​the renal sinus and the surrounding renal parenchyma, as well as the radial arrangement of the renal cortex and medulla. The coronal section is parallel to the x-axis and z-axis of the virtual coordinate system, with the normal direction being the y-axis of the virtual coordinate system. The coronal section shows the outline of the lateral edge of the kidney and the boundary between the renal parenchyma and the perirenal fat.

[0009] Furthermore, in the definition of the intelligent agent state, a three-dimensional coordinate system is used to represent the position of the virtual ultrasound probe; the position of the virtual ultrasound probe is determined by the three coordinate values ​​of x, y, and z; the x, y, and z values ​​reflect the specific position of the virtual ultrasound probe in the space around the kidney; the Euler angles (α, β, γ) are used to describe the rotation angles of the probe around the x, y, and z axes; among them, α represents the rotation angle around the x-axis of the virtual coordinate system, β represents the rotation angle around the y-axis of the virtual coordinate system, and γ represents the rotation angle around the z-axis of the virtual coordinate system.

[0010] Further, the virtual ultrasound probe can perform discrete actions including translational actions and rotational actions. Among them, the translational actions include translation along the x-axis and translation along the y-axis, and the rotational actions include rotation around the x-axis, rotation around the y-axis, and rotation around the z-axis. The translation and rotation step sizes are dynamically adjusted according to the current reward value.

[0011] Further, the ultrasound images of several recent time steps are input into a pre-trained UNet network, and the extracted feature vectors include: The pre-trained UNet network is trained based on medical image data. Its encoder part extracts the spatial features in the image layer by layer through multi-layer convolution and pooling operations, and the decoder part fuses the features of different levels to retain the detailed information; a sliding window is maintained, and the ultrasound images within the sliding window are sequentially input into the UNet network. The feature map output by the last layer of the encoder is subjected to global average pooling to obtain a feature vector containing the spatial context and the dynamic changes in the time dimension; this feature vector is concatenated with the probe pose information at the current time step to jointly constitute the input state of the DQN model.

[0012] Further, the DQN model is trained, specifically including: initializing the network parameters, setting the number of training steps as N, and the number of steps currently in the training as n. During the training process, it is judged whether n is less than N. If n = N, the training ends and QNet is output. If n < N, the agent is reset, and the time, step size, angle, and random pose of the agent are initialized. According to the ultrasound image of the current pose, the feature vector is calculated through the UNet network and used as the current state. An action is selected according to the greedy policy, the pose is updated, and the new state and reward are calculated. The previous state, the executed action, the reward, and the current state are stored in the experience replay pool. The step size is adjusted according to the situation, and the number of steps is updated to make n = n + 1. It is judged whether n is a multiple of i. If n is a multiple of i, QNet is updated. If n is not a multiple of i, it is judged whether n is a multiple of k. If n is a multiple of k, the parameters of the QNet network are copied to the target network and then directly enter the next loop, and it is judged whether n is less than N; if n is not a multiple of k, it directly enters the next loop and judges whether n is less than N. The above process is repeated until n = N, and then the training ends and QNet is output; i represents the update interval steps of the main network QNet, and k represents the interval steps for the target network parameters to synchronize with the main network parameters.

[0013] Furthermore, the trained DQN model is used to control the movement of the actual ultrasound probe. Specifically, the UNet network receives the current ultrasound image from the actual ultrasound probe at every time step and extracts the image's feature vector. The extracted feature vector is used as the input state of the DQN model. The DQN model calculates the Q value of each possible action based on the input state and selects the action with the largest Q value. This action is the optimal posture adjustment solution for the probe.

[0014] The present invention proposes a deep reinforcement learning-based autonomous renal ultrasound navigation method. By training a DQN model in a virtual simulation environment, the ultrasound probe autonomously learns a hierarchical search strategy from coarse positioning to fine adjustment. This technology dynamically optimizes action decisions to approximate the standard imaging plane by combining ultrasound image spatiotemporal features with probe pose information in real time. This technology significantly reduces physician manual operation time, lowers the barrier to entry, and improves the efficiency and stability of renal ultrasound examinations, possessing significant clinical application value. The reward function deeply integrates ultrasound image confidence (such as renal sinus clarity and cortical-medullary boundary features extracted via a UNet network) with probe pose errors (distance and pose deviation). When the hyperechoic renal sinus area is fully displayed and the cortical-medullary layer is clearly defined in the image, the agent receives a positive reward, driving the probe to optimize toward that pose. If the image edges are blurred or the renal structure is incompletely displayed, a negative penalty is applied, triggering corrective actions. This mechanism evaluates image quality in real time during navigation and dynamically adjusts the probe pose, increasing the success rate of obtaining qualified images. This effectively reduces repeated scans due to insufficient image quality and provides a stable imaging foundation for accurate diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flowchart of the deep reinforcement learning-based autonomous renal ultrasound navigation method of this application; Figure 2 Flowchart for training the DQN model for this application. DETAILED DESCRIPTION

[0016] This embodiment of the present invention discloses a method for autonomous renal ultrasound navigation based on deep reinforcement learning. This method applies the Deep Q-Network (DQN) algorithm to address the position and image recognition issues in autonomous renal ultrasound imaging navigation. The DQN algorithm combines deep neural networks with Q-learning technology. DQN uses deep neural networks to approximate the Q-value function, which represents the long-term cumulative reward obtained by taking an action in a given state. In DQN, the Q-learning update rule is as follows: ; in is the status Take action Q value, It’s an instant reward. is the learning rate, is a discount factor that balances the weight of current rewards and future rewards. Through this formula, DQN can select the optimal strategy that can obtain the maximum reward in different states.

[0017] This application discloses a method for autonomous kidney ultrasound navigation based on deep reinforcement learning, such as Figure 1 , including the steps of: S1. Use 3D voxel point clouds to construct a 3D virtual environment containing a kidney, construct an intelligent agent to control a virtual ultrasound probe in the 3D virtual environment, configure the kidney to the region of interest in the 3D virtual environment, and set the target standard imaging plane. S2. Define the state of the intelligent agent as consisting of the position of the virtual ultrasound probe and the ultrasound image features, and define that the virtual ultrasound probe can perform several discrete actions. S3. Prepare a pre-trained UNet network and a DQN model, collect the current position and ultrasound image of the virtual ultrasound probe at each time step, input the ultrasound images of the most recent several time steps into the pre-trained UNet network, extract the feature vector, and use it as the input state of the DQN model. S4. Train the DQN model. During the training process, set a reward function, configure the intelligent agent to receive higher rewards the closer it is to the target standard imaging plane, and configure the intelligent agent to be penalized when it leaves the region of interest or when an edge kidney is displayed on the ultrasound image. Store historical state, action, and reward data in the experience replay pool, and regularly synchronize the main network parameters with the target network for training, ultimately forming a DQN model that can autonomously locate the standard imaging plane. Model; S5. Use the trained DQN model to control the movement of the actual ultrasound probe.

[0018] For S1. Use the three-dimensional voxel point cloud to construct a three-dimensional virtual environment of the phantom containing the kidney, build an intelligent body to control the virtual ultrasound probe in the three-dimensional virtual environment of the phantom, configure the kidney to the region of interest in the three-dimensional virtual environment of the phantom, and set the target standard imaging plane.

[0019] Specifically, the 3D virtual environment of the kidney phantom is constructed using medical imaging data, such as cross-sectional data from real kidney CT / MRI scans, to construct a layered voxel point cloud model of anatomical structures such as the cortex, medulla, and renal sinus. Within this 3D virtual environment, the kidney is fixed at the origin of the virtual coordinate system and surrounded by a biomimetic tissue medium. The acoustic properties of the biomimetic tissue medium match those of human soft tissue, simulating the propagation and attenuation of ultrasound waves and generating dynamic ultrasound echo signals.

[0020] Specifically, an intelligent agent was constructed within a three-dimensional virtual environment of a body model. The agent possesses perception, decision-making, and memory functions. Specifically, the agent comprises a perception unit, a decision-making unit, and a memory unit. The perception unit acquires the spatial pose information of the virtual ultrasound probe in real time, including its coordinate position along the x, y, and z axes and its rotation angles around each axis. It also collects ultrasound image data at the corresponding position of the virtual ultrasound probe. The decision-making unit integrates a DQN model, taking the pose output by the perception unit and ultrasound image features as input states. It selects the optimal action by calculating the Q value of each discrete action (translation along the x and y axes, rotation along the x, y, and z axes, a total of 10 operations) and drives the virtual ultrasound probe to perform the corresponding pose adjustment. The memory unit is responsible for storing historical state, action, and reward data in an experience replay pool, providing data support for DQN model training and optimization.

[0021] It's understandable that through the interactive loop of perception, decision-making, execution, and memory, the intelligent agent continuously learns in the virtual environment, gradually mastering strategies for efficient navigation to the target standard imaging plane. This ensures unified coordination with the overall solution's reinforcement learning training, action space setting, and data storage. The virtual ultrasound probe simulates the physical properties and motion mechanisms of a real ultrasound probe within the phantom's 3D virtual environment, capable of translation along the x- and y-axes and rotation around them. It can perform a total of 10 discrete actions, including positive and negative translations along the x- and y-axes, and positive and negative rotations along the x-, y-, and z-axes (where the z-axis translation is determined by the x- and y-axis translations combined with the soft body surface model). The step sizes of these actions follow a coarse-to-fine hierarchical strategy: Initially, large translation and rotation step sizes are used to facilitate rapid adjustment of the probe's position over a wide range, allowing the agent to quickly approach the approximate area of ​​the target standard imaging plane and shortening the initial search time. As training progresses, when the agent approaches the target area, the step size is automatically reduced, and fine adjustments are made to ensure that the virtual ultrasound probe can accurately align with the target standard imaging plane, achieving a transition from rapid coarse positioning to precise fine positioning. Furthermore, the virtual ultrasound probe provides real-time feedback of its position at each time step, which, together with ultrasound image features, constitutes the agent's input state, providing the DQN model with accurate environmental perception data to generate optimal action decisions, drive probe position adjustments, and continuously optimize towards the target standard imaging plane.

[0022] Specifically, the region of interest is set as a cubic space containing the complete anatomical structure of the kidney, covering tissues such as the renal cortex, medulla, and renal sinus. The specific boundaries are adjusted according to the average kidney size to ensure that the movement of the virtual ultrasound probe in this area can effectively cover the posture range required for standard kidney imaging.

[0023] Specifically, the target standard imaging planes are based on typical standard sections used in clinical renal ultrasound examinations. These include the longitudinal, transverse, and coronal planes, each of which is strictly aligned with the virtual coordinate system. The longitudinal plane is parallel to the long axis of the kidney (z-axis) and the anteroposterior direction (y-axis), with its normal oriented to the x-axis. It must fully display the long-axis anatomy of the kidney, including the hilum, sinus, cortex, and medulla. The transverse plane is perpendicular to the long axis of the kidney (z-axis), parallel to the xy plane, with its normal oriented to the z-axis. It must present a short-axis cross-section of the kidney, clearly demonstrating the circumferential distribution of the hyperechoic sinus and surrounding renal parenchyma, as well as the radial arrangement of the renal cortex and medulla. The coronal plane is parallel to the short axis (x-axis) and long axis (z-axis), with its normal oriented to the y-axis. It must demonstrate the outline of the kidney's lateral margin and the boundary between the renal parenchyma and perirenal fat, allowing for assessment of the kidney's overall morphology and adjacent relationships.

[0024] For S2. Define the state of the intelligent agent as consisting of the posture of the virtual ultrasound probe and the ultrasound image features, and define that the virtual ultrasound probe can perform several discrete actions.

[0025] Specifically, in the agent state definition, a three-dimensional coordinate system is used to accurately represent the position of the virtual ultrasound probe. The position of the virtual ultrasound probe is determined by three coordinate values: x, y, and z. The x, y, and z values ​​reflect the specific position of the virtual ultrasound probe in the space around the kidney, allowing the agent to clearly understand the spatial position relationship of the virtual ultrasound probe relative to the kidney. Euler angles ( α , β , γ ) is used to describe the rotation angle of the probe around the x, y, and z axes. α represents the rotation angle around the x-axis, β represents the rotation angle around the y-axis, γ Represents the rotation angle around the z-axis. This angle information accurately defines the probe's orientation, which is crucial for acquiring ultrasound images at different angles.

[0026] Specifically, the virtual ultrasound probe can perform discrete actions including translation and rotation, where the translation action includes translation along the x-axis and translation along the y-axis. Two discrete actions are defined along the x-axis, namely positive translation and negative translation. In the early stages of training, the translation step size is set large, such as ±10 mm each time, so that the agent can quickly adjust the probe position within a wider range. As the training progresses, when the agent approaches the target area, the step size gradually decreases, such as to ±1 mm, to achieve fine adjustment. Translation along the y-axis also defines two translation actions, positive and negative, and the step size adjustment strategy is the same as that of translation along the x-axis. This translation operation allows the probe to move in the front-to-back direction, thereby acquiring ultrasound images at different positions.

[0027] Specifically, the rotation actions include rotation around the x-axis, rotation around the y-axis, and rotation around the z-axis. The rotation around the x-axis has two discrete actions: positive rotation and negative rotation. The rotation around the y-axis includes positive and negative rotation actions. The step size adjustment method is the same as that of the rotation around the x-axis. The rotation around the z-axis also has two actions: positive and negative rotation.

[0028] Understandably, at the beginning of training, the agent has little knowledge of the location of the target standard imaging plane, so a larger action step size is used. This allows the probe to cover a larger spatial range in a shorter time and quickly approach the target area. For example, when searching for the approximate location of the kidney, the probe can quickly adjust to the area that may contain the target plane through translation and rotation movements with larger steps. When the agent determines that the probe is close to the target standard imaging plane based on the reward function, the action step size is automatically reduced. A smaller step size allows the probe to make more precise adjustments, thereby accurately aligning it with the target plane. For example, after the location of the target plane has been roughly determined, the probe can be gradually and accurately positioned on the standard imaging plane through translation and rotation movements with small steps.

[0029] For S3. Prepare the pre-trained UNet network and DQN model, collect the current position of the virtual ultrasound probe and the ultrasound image at each time step, input the ultrasound images of the most recent several time steps into the pre-trained UNet network, extract the feature vector, and use it as the input state of the DQN model.

[0030] Specifically, the current position of the virtual ultrasound probe and ultrasound image are captured at each time step. The ultrasound images from the most recent several time steps are fed into a pre-trained UNet network to extract feature vectors. Within the virtual simulation environment, the agent's perception unit captures the virtual ultrasound probe's 3D spatial pose (including its x, y, and z coordinates and rotation angles around each axis) in real time at each time step, simultaneously acquiring the ultrasound image at that pose. A sliding window of length 4 is maintained, storing the ultrasound images from the most recent four time steps in chronological order, forming an image group containing time series information.

[0031] The pre-trained UNet network is trained based on medical image data. In its encoder part, through multi-layer convolution and pooling operations, spatial features in the image are extracted layer by layer (such as kidney edges, corticomedullary junction, renal sinus structure, etc.). The decoder part then fuses features at different levels to retain detailed information. The 4-frame ultrasound images within the sliding window are sequentially input into the UNet network. The feature map output by the last layer of the encoder is subjected to global average pooling to obtain a feature vector containing spatial context and dynamic changes in the time dimension. This feature vector is concatenated with the probe pose information (numerical representation of position coordinates and rotation angle) at the current time step to jointly form the input state of the DQN model, enabling the agent to simultaneously perceive the spatial position and pose of the probe and the spatio-temporal features of the ultrasound image, providing multi-dimensional environmental state information for making optimal decisions.

[0032] For S4. Training the DQN model. During the training process, set the reward function. Configure that the closer the agent is to the target standard imaging plane, the higher the reward it obtains. Configure that when the agent leaves the region of interest or the edge of the kidney appears in the ultrasound image, it is penalized. Store the historical state, action, and reward data in the experience replay pool, and regularly synchronize the parameters of the main network to the target network for training. Eventually, a DQN model capable of autonomously locating the standard imaging plane is formed.

[0033] Specifically, in S4, the DQN model is trained as Figure 2 , specifically including the steps: Initialize the network parameters. Set the number of training steps to N, and the current number of training steps to n. During the training process, judge whether n is less than N. If n = N, end the training and output QNet. If n < N, reset the agent, initialize the time, step size, angle, and random pose of the agent. According to the ultrasound image of the current pose, calculate the feature vector through the UNet network, and use it as the current state. Select an action according to the greedy strategy, update the pose, and calculate the new state and reward. Store the previous state, executed action, reward, and current state in the experience replay pool. Adjust the step size according to the situation, and update the number of steps to make n = n + 1. Judge whether n is a multiple of i. If n is a multiple of i, update QNet. If n is not a multiple of i, judge whether n is a multiple of k. If n is a multiple of k, copy the parameters of the QNet network to the target network and then directly enter the next loop to judge whether n is less than N; if n is not a multiple of k, directly enter the next loop to judge whether n is less than N. Repeat the above process until n = N, then end the training and output QNet; i represents the update interval of the main network QNet. Every time the training step number n is a multiple of i, the parameters of QNet are updated so that it can optimize the Q value according to the new experience data in time and adjust the network to learn the optimal strategy.

[0034] k represents the interval number of steps for synchronizing the target network parameters with the main network (QNet) parameters. When the number of steps n is a multiple of k, the parameters of QNet are copied to the target network, so that the target network can obtain updates from the main network while maintaining relative stability, which is used to calculate a stable target Q value, reduce fluctuations during training, and improve training stability.

[0035] For S5, the trained DQN model is used to control the movement of the actual ultrasound probe.

[0036] In S5, the trained DQN model is used to control the movement of the actual ultrasound probe. Specifically, the UNet network receives the current ultrasound image from the actual ultrasound probe at every time step and extracts the image's feature vector. The extracted feature vector is used as the input state of the DQN model. The DQN model calculates the Q value of each possible action based on the input state and selects the action with the largest Q value. This action is the optimal posture adjustment solution for the probe. The specific details are as follows: The trained DQN model consists of a main network (QNet) and a target network. The main network receives a composite input state composed of probe pose information and ultrasound image spatiotemporal features. The main network nonlinearly maps these composite features through three fully connected layers, outputting Q values ​​corresponding to 10 discrete actions. These Q values ​​represent the long-term cumulative reward the agent can obtain by performing an action in the current state. This calculation incorporates a weighted evaluation of a multi-dimensional reward function: Positive reward calculation: Based on the shortest Euclidean distance between the current position of the probe and the target plane, a reward of +3 is given for every 1mm reduction in distance. When the distance is less than 5mm, a saturation reward of +15 is triggered. Calculate the Euler angle deviation (the sum of the absolute values ​​of the α / β / γ angles) between the probe normal and the target plane normal. Each 1° decrease in the Euler angle will earn you a +2 reward. If the deviation is less than 3°, you will receive a +10 reward. The feature activation values ​​of the renal sinus hyperechoic area and the cortex-medullary boundary in the UNet feature vector are mapped to a reward of 0-10 points (the higher the clarity, the greater the activation value, and the higher the reward).

[0037] Negative penalty mechanism: If the probe goes beyond the region of interest (cube space boundary), a penalty of -20 is immediately imposed and a reset is triggered; If the kidney edge gradient in the ultrasound image is lower than the preset threshold (edge ​​blur determined by UNet features), a penalty of -10 is applied and the Q value of the current action direction is suppressed.

[0038] The optimal action command output by the DQN model is transmitted to the ultrasound probe driver module via a real-time control interface, driving the robotic arm or servo motor to adjust the probe's position. After the adjustment, the system immediately captures ultrasound images and position data for the new state, calculates an immediate reward, and stores the "state-action-reward-new-state" data in an experience replay pool, which is used to periodically update the main network parameters (every i = 100 steps) and synchronize the target network (every k = 1000 steps). Through a closed-loop control process of "perception-decision-execution-feedback," the intelligent agent continuously optimizes its Q-value strategy, ensuring that the probe gradually approaches and stabilizes on the target standard imaging plane during dynamic adjustments, achieving autonomous navigation without human intervention.

[0039] It is obvious that the method of the present invention can be implemented by a computer program. The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0040] Therefore, it can be understood that the present invention discloses an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above method.

[0041] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

[0042] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the scope of protection of the present invention.

Claims

1. A deep reinforcement learning-based autonomous ultrasound navigation method for the kidney, characterized in that: Including steps: A 3D virtual environment containing the kidney is constructed using a 3D voxel point cloud. An intelligent agent is constructed in the 3D virtual environment to control a virtual ultrasound probe. The kidney is positioned within the 3D virtual environment to locate the region of interest, and a target standard imaging plane is set. The state of the agent is defined to consist of the posture of the virtual ultrasound probe and the ultrasound image features, and the virtual ultrasound probe is defined to be able to perform several discrete actions; Prepare a pre-trained UNet network and DQN model. Collect the current position of the virtual ultrasound probe and ultrasound image at each time step. Input the ultrasound images of the most recent several time steps into the pre-trained UNet network and extract the feature vectors, which serve as the input state of the DQN model. Train the DQN model, store historical states, actions, and reward data in the experience replay pool, and regularly synchronize the main network parameters with the target network for training, ultimately forming a DQN model that can autonomously locate the standard imaging plane; The trained DQN model is used to control the movement of the actual ultrasound probe.

2. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: A three-dimensional virtual environment of the kidney phantom is constructed using medical imaging data, constructing a layered voxel point cloud model that includes the anatomical structures of the cortex, medulla, and renal sinus. In this three-dimensional virtual environment, the kidney is fixed at the origin of the virtual coordinate system and is surrounded by a bionic tissue medium. The acoustic properties of the bionic tissue medium match those of human soft tissue, simulating the propagation and attenuation characteristics of ultrasound to generate dynamic ultrasound echo signals.

3. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: The intelligent agent includes a perception unit, a decision unit and a memory unit. The perception unit is used to obtain the spatial posture information of the virtual ultrasound probe in real time and collect ultrasound image data of the corresponding position of the virtual ultrasound probe; the decision unit integrates the DQN model. The decision unit uses the posture output by the perception unit and the ultrasound image features as input states, calculates the state-action value function Q value of each discrete action to select the optimal action; the memory unit is used to store historical states, actions, and reward data in the experience replay pool to provide data support for the training optimization of the DQN model.

4. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: The region of interest is set as a cubic space containing the complete anatomical structure of the kidney, and the region of interest includes the renal cortex, medulla, and renal sinus tissue.

5. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: The target standard imaging planes are set based on typical standard sections used in clinical renal ultrasound examinations. These target standard imaging planes include three core planes: longitudinal, transverse, and coronal. Each plane definition is aligned with the virtual coordinate system. The longitudinal section is parallel to the z-axis and y-axis of the virtual coordinate system, with its normal direction being the x-axis of the virtual coordinate system. The longitudinal section fully displays the long-axis anatomical structure of the kidney, including the hilum, renal sinus, cortex, and medulla. The transverse section is perpendicular to the z-axis of the virtual coordinate system, parallel to the xy plane of the virtual coordinate system, and with its normal direction being the z-axis of the virtual coordinate system. The transverse section presents a short-axis cross-section of the kidney, showing the annular distribution of the renal sinus hyperechoic area and the surrounding renal parenchyma, as well as the radial arrangement of the renal cortex and medulla. The coronal section is parallel to the x-axis and z-axis of the virtual coordinate system, with its normal direction being the y-axis of the virtual coordinate system. The coronal section shows the outline of the lateral edge of the kidney and the boundary between the renal parenchyma and perirenal fat.

6. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: In the definition of the agent state, a three-dimensional coordinate system is used to represent the position of the virtual ultrasound probe; the position of the virtual ultrasound probe is determined by three coordinate values: x, y, and z; the x, y, and z values ​​reflect the specific position of the virtual ultrasound probe in the space around the kidney; the Euler angle ( α , β , γ ) describes the rotation angle of the probe around the x, y, and z axes; where, α represents the rotation angle around the x-axis of the virtual coordinate system, β Indicates the rotation angle around the y-axis of the virtual coordinate system, γ Indicates the rotation angle around the z-axis of the virtual coordinate system.

7. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: The virtual ultrasound probe can perform discrete actions including translational actions and rotational actions. Among them, the translational actions include translation along the x-axis and translation along the y-axis, and the rotational actions include rotation around the x-axis, rotation around the y-axis, and rotation around the z-axis. The translation and rotation step sizes are dynamically adjusted according to the current reward value.

8. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: Input the ultrasound images of several recent time steps into the pre-trained UNet network, and the extracted feature vectors include: The pre-trained UNet network is trained based on medical image data. Its encoder part extracts the spatial features in the image layer by layer through multi-layer convolution and pooling operations, and the decoder part fuses the features of different levels to retain the detailed information; maintain a sliding window, input the ultrasound images within the sliding window into the UNet network in sequence, take the feature map output by the last layer of the encoder for global average pooling, and obtain a feature vector containing the spatial context and the dynamic changes in the time dimension; this feature vector is concatenated with the probe pose information at the current time step to jointly form the input state of the DQN model.

9. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: Train the DQN model, specifically including: initialize the network parameters, set the number of training steps as N, the current number of steps in the training as n, during the training process, judge whether n is less than N. If n = N, then end the training and output QNet. If n < N, then reset the agent, initialize the time, step size, angle and random pose of the agent, calculate the feature vector through the UNet network according to the ultrasound image of the current pose, use it as the current state, select an action according to the greedy strategy, update the pose, and calculate the new state and reward. Store the previous state, executed action, reward and current state in the experience replay pool, adjust the step size according to the situation, update the number of steps to make n = n + 1, judge whether n is a multiple of i. If n is a multiple of i, then update QNet. If n is not a multiple of i, then judge whether n is a multiple of k. If n is a multiple of k, then copy the parameters of the QNet network to the target network and then directly enter the next loop, judge whether n is less than N; if n is not a multiple of k, then directly enter the next loop, judge whether n is less than N, repeat the above process until n = N, then end the training and output QNet; i represents the update interval number of steps of the main network QNet, and k represents the interval number of steps for the target network parameters to synchronize the main network parameters.

10. The method for autonomous renal ultrasound navigation based on deep reinforcement learning according to claim 1, characterized in that: Use the trained DQN model to control the movement of the actual ultrasound probe. Specifically, the UNet network receives the current ultrasound image from the actual ultrasound probe every other time step and extracts the feature vector of the image. The extracted feature vector is used as the input state of the DQN model. The DQN model calculates the Q values of each possible action according to the input state and selects the action with the largest Q value. This action is the optimal pose adjustment plan for the probe.

Citation Information

Patent Citations

  • Method and system for carrying out ultrasonic automatic detection based on ultrasonic image

    CN116327241A