Path planning method and device applied to magnetic drive micro-robot and electronic equipment

By acquiring state and fluid information, using Transformer modules and multi-layer perceptrons to extract features, and combining reward functions with reinforcement learning strategy networks for path planning, the problem of inaccurate path planning of magnetically driven spiral robots in complex dynamic fluid environments is solved, achieving safe and efficient path planning.

CN120609359APending Publication Date: 2025-09-09SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510763888.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

The path planning of the magnetically driven spiral robot is not accurate enough in complex dynamic fluid environments, resulting in low motion efficiency and high energy consumption, which cannot meet clinical needs.

Method used

By acquiring state information and fluid information, using the Transformer module and multi-layer perceptron to extract features, combining the reward function and reinforcement learning strategy network for path planning, a reward mechanism is constructed to optimize the path of the magnetically driven microrobot to ensure that it reaches the target location safely and efficiently.

Benefits of technology

Safe and energy-saving path planning of magnetically driven microrobots is achieved in a dynamic flow field environment, which improves the accuracy and efficiency of path planning. It is suitable for scenarios such as medical microrobots with strict requirements on reliability and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120609359A_ABST
    Figure CN120609359A_ABST
Patent Text Reader

Abstract

The invention provides a path planning method and device applied to a magnetic drive micro-robot, electronic equipment and a storage medium, and relates to the technical field of robots. The method comprises the following steps: acquiring state information and fluid information; on the basis of the state information and the fluid information, capturing the moving characteristics of the magnetic drive micro-robot in the dynamic flow field environment with the obstacle, and obtaining target characteristics; according to a reward function constructed by energy consumed by the magnetic drive micro-robot in the moving process of the dynamic flow field environment with the obstacles and the target characteristics, path planning is conducted on the next action of the magnetic drive micro-robot, and path information is obtained; and according to the path information, controlling the magnetic drive micro-robot to move in the dynamic flow field environment with the obstacle, so that the magnetic drive micro-robot approaches or reaches a target position. The problem that the path planning of the magnetic drive spiral robot is not accurate enough in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robotics technology. Specifically, the present application relates to a path planning method, device, electronic device and storage medium for a magnetically driven micro robot. Background Art

[0002] Traditional oral or intravenous drug delivery methods only achieve accurate delivery to the lesion site in 0.7%, resulting in very low delivery efficiency. Magnetic fields are safe and radiation-free, and have been widely used in the biomedical field. Helical microrobots are considered an ideal model for drug transport in low-Reynolds-number biofluid environments.

[0003] Drug delivery requires precise delivery of drugs to specific targets within the body to enhance therapeutic efficacy and minimize damage to healthy tissue. Furthermore, robots must be able to navigate the confined and complex environment of the body, avoiding vital organs and blood vessels to safely and efficiently reach the site of disease. Traditional surgical methods struggle to achieve such high-precision operations at such a microscopic scale. The magnetically driven helical robot, with its miniature size and flexible maneuverability, is poised to fill this gap.

[0004] In the complex and dynamic in vivo environment, magnetically driven helical robots have significant limitations. Most products have simple control algorithms that fail to fully consider the complex dynamic environment. Path planning struggles with time-varying fluid dynamics, resulting in low efficiency and high energy consumption in dynamic fluids. These limitations fail to meet the clinical demands for efficient and safe magnetically driven helical robots.

[0005] From the above, we can see that the problem of inaccurate path planning of magnetically driven spiral robots still needs to be solved. Summary of the Invention

[0006] This application provides a path planning method, device, electronic device, and storage medium, which can solve the problem of inaccurate path planning of magnetically driven spiral robots in related technologies. The technical solution is as follows:

[0007] According to one aspect of the present application, a path planning method applied to a magnetically driven microrobot includes: obtaining state information and fluid information; the state information is used to describe the spatial state of the magnetically driven microrobot and an obstacle; the fluid information is used to describe the dynamic flow field environment in which the magnetically driven microrobot is located; based on the state information and the fluid information, the movement characteristics of the magnetically driven microrobot in the dynamic flow field environment where an obstacle is present are captured to obtain target characteristics; according to a reward function constructed based on the energy consumed by the magnetically driven microrobot during the movement in the dynamic flow field environment where an obstacle is present, and the target characteristics, path planning is performed for the next action of the magnetically driven microrobot to obtain path information; according to the path information, the magnetically driven microrobot is controlled to move in the dynamic flow field environment where an obstacle is present, so that the magnetically driven microrobot approaches or reaches the target position.

[0008] According to one aspect of the present application, a path planning device for a magnetically driven microrobot includes: an information acquisition module for acquiring state information and fluid information; the state information is used to describe the spatial state of the magnetically driven microrobot and an obstacle; the fluid information is used to describe the dynamic flow field environment in which the magnetically driven microrobot is located; a feature processing module for capturing the movement characteristics of the magnetically driven microrobot in the dynamic flow field environment where obstacles exist based on the state information and the fluid information, and obtaining target characteristics; a path planning module for performing path planning for the next action of the magnetically driven microrobot based on a reward function constructed according to the energy consumed by the magnetically driven microrobot during the movement in the dynamic flow field environment where obstacles exist, and the target characteristics, and obtaining path information; a motion control module for controlling the magnetically driven microrobot to move to a target position in the dynamic flow field environment where obstacles exist based on the path information.

[0009] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the path planning method described above is implemented.

[0010] According to one aspect of the present application, a storage medium stores a computer program thereon, and when the computer program is executed by one or more processors, the path planning method described above is implemented.

[0011] According to one aspect of the present application, a computer program product includes a computer program, and when the computer program is executed by one or more processors, the computer program implements the path planning method described above.

[0012] The beneficial effects of the technical solution provided by this application are:

[0013] In the above technical solution, the reward function constructed based on the energy consumed by the magnetically driven microrobot during its movement in a dynamic flow field environment with obstacles, as well as the target characteristics, is used to plan the path for the next action of the magnetically driven microrobot. The fluid velocity field, obstacle distribution and energy consumption are taken into consideration. The direction and intensity of the magnetic field are adjusted in real time through the magnetic field control system, so that the magnetically driven microrobot can avoid obstacles along the path, overcome fluid resistance, reduce energy consumption, and move to the target position accurately, ensuring the completion and safety of the task. This can effectively solve the problem of inaccurate path planning of magnetically driven spiral robots in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts.

[0015] Figure 1 It is a schematic diagram of the implementation environment involved in this application;

[0016] Figure 2 is a hardware structure diagram of an electronic device according to an exemplary embodiment;

[0017] Figure 3 is a flow chart showing a path planning method according to an exemplary embodiment;

[0018] Figure 4 yes Figure 3 A flowchart of step 330 in one embodiment corresponding to the embodiment;

[0019] Figure 5 yes Figure 3 The steps before step 350 in the corresponding embodiment are a flowchart of an embodiment;

[0020] Figure 6 yes Figure 3 A flowchart of an embodiment corresponding to step 350 in an embodiment;

[0021] Figure 7 yes Figure 6 The steps following step 353 in the corresponding embodiment are a flowchart of an embodiment;

[0022] Figures 8a to 8c This is a schematic diagram of a specific implementation of a path planning method in an application scenario;

[0023] Figure 9is a structural block diagram of a path planning device according to an exemplary embodiment;

[0024] Figure 10 The figure is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0025] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0026] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0027] As mentioned above, magnetically driven helical robots have obvious limitations in actual complex dynamic in vivo environments.

[0028] Some magnetically driven helical robots are already capable of basic motion under the control of a stable magnetic field in specific laboratory environments or simple simulated in vivo environments, demonstrating a certain degree of control. However, in dynamic in vivo fluid environments (such as blood flow and tissue fluid flow), magnetically driven helical robots face numerous challenges, such as the time-varying characteristics of dynamic fluids, limited energy supply within a specific timeframe, and the risk of collision with surrounding tissue or other foreign objects. These challenges hinder their effective application in complex in vivo environments.

[0029] Currently, various path planning methods attempt to address the path planning problem for magnetically driven helical robots, such as particle swarm optimization (PSO) and rapidly exploring random trees (RRT). However, these methods primarily focus on path optimization in static environments and fail to consider time-varying fluid dynamics for real-time planning. Numerical optimization techniques are computationally expensive and unsuitable for real-time planning of magnetically driven microrobots. Local planning methods, such as artificial potential fields, lack convergence guarantees in complex flow fields and with interacting obstacles. Genetic algorithm-based methods suffer from slow convergence, low computational efficiency, and a tendency to obtain suboptimal solutions. Fuzzy logic hybrid methods face challenges in handling high-density dynamic obstacles due to complex rules. Methods based on random waypoints generate non-smooth paths, which affects real-time directional adaptability. Among learning-based methods, some reinforcement learning frameworks suffer from low data efficiency and difficulty in real-time convergence, resulting in suboptimal paths planned in dynamic environments. Furthermore, existing research has not addressed the energy-efficient path planning problem for magnetically driven microrobots in complex, time-varying flow fields.

[0030] From the above, it can be seen that the relevant technologies cannot effectively solve the energy-saving path planning problem of the magnetically driven spiral robot in a complex fluid environment, which limits the application of the magnetically driven spiral robot.

[0031] To this end, the path planning method provided in this application can effectively improve the inaccurate path planning of the magnetically driven spiral robot. Accordingly, the path planning method is suitable for a path planning device, which can be deployed in an electronic device. The electronic device can be a computer device configured with a von Neumann architecture, for example, the computer device includes a desktop computer, a laptop computer, a server, etc.

[0032] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0033] Figure 1 It should be noted that this implementation environment is only an example adapted to the present invention and should not be considered as providing any limitation on the scope of application of the present invention.

[0034] The implementation environment includes a collection end 110 and a service end 130 .

[0035] Specifically, the acquisition end 110 may be a magnetic sensing and magnetic control feedback system.

[0036] Server 130 can be an electronic device such as a desktop computer, laptop computer, or server, or a computer cluster consisting of multiple servers, or even a cloud computing center consisting of multiple servers. Server 130 is used to provide background services, such as, but not limited to, path planning services.

[0037] A network communication connection is pre-established between the server 130 and the collection terminal 110 via a wired or wireless method, and data transmission between the server 130 and the collection terminal 110 is achieved through the network communication connection. The transmitted data includes but is not limited to: state information and fluid information.

[0038] In an application scenario, through the interaction between the collection terminal 110 and the service terminal 130 , the collection terminal 110 collects state information corresponding to the target and a dynamic flow field environment corresponding to the dynamic flow field environment.

[0039] For the server 130, after receiving the status information and fluid information uploaded by the acquisition terminal 110, it calls the path planning service, performs feature extraction and feature fusion processing based on the status information and fluid information, and obtains the target features; according to the constructed reward function, the path planning is performed for the next action of the target based on the target features to obtain the path information; according to the path information, the target is controlled to move in the dynamic flow field environment, so as to solve the problem of inaccurate path planning of the magnetically driven spiral robot in related technologies.

[0040] See also Figure 2 , Figure 2 This is a hardware structure diagram of an electronic device according to an exemplary embodiment. Figure 1 A server 130 is shown in an implementation environment.

[0041] It should be noted that the electronic device is only an example adapted for this application and cannot be considered to provide any limitation on the scope of use of this application. The electronic device cannot be interpreted as needing to rely on or must have Figure 2 One or more components of exemplary electronic device 200 are shown.

[0042] The hardware structure of the electronic device 200 may vary greatly due to different configurations or performances, such as Figure 2 As shown, the electronic device 200 includes a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .

[0043] Specifically, the power supply 210 is used to provide operating voltage for various hardware devices on the electronic device 200 .

[0044] The interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. Figure 1 The interaction between the collection end 110 and the service end 130 in the implementation environment is shown.

[0045] Of course, in other examples adapted by this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input-output interface 235, and at least one USB interface 237, etc. Figure 2 As shown, this does not constitute a specific limitation.

[0046] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include an operating system 251, application 253 and data 255, etc. The storage method can be temporary storage or permanent storage.

[0047] Among them, the operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 200 to enable the central processing unit 270 to calculate and process the massive data 255 in the memory 250. It can be WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0048] The application program 253 is a computer program formed by computer-readable instructions based on the operating system 251 to perform at least one specific task, and may include at least one module ( Figure 2 Each module may include corresponding computer-readable instructions. For example, the path planning device may be considered as an application 253 deployed on the electronic device 200.

[0049] The data 255 may be photos, pictures, etc. stored in a disk, or may be status information, fluid information, etc. stored in the memory 250 .

[0050] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read the computer program stored in the memory 250, thereby performing operations and processing on the massive data 255 in the memory 250. For example, the path planning method can be implemented by the central processing unit 270 reading the application program 253 stored in the memory 250.

[0051] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.

[0052] See also Figure 3 , an embodiment of the present application provides a path planning method, which is applicable to an electronic device, for example, the electronic device may be Figure 1 The server 130 in the implementation environment is shown, and the hardware structure of the electronic device can be as follows Figure 2 shown.

[0053] In the following method embodiments, for ease of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.

[0054] like Figure 3 As shown, the method may include the following steps:

[0055] Step 310: Acquire status information and fluid information.

[0056] First of all, it should be noted that the magnetically driven microrobot is a small robot that can be driven and controlled by an external magnetic field. It is soft and compliant and can enter complex and narrow areas. It has a spiral structure and can perform controllable rotational motion under the action of a magnetic field. It can adapt to complex environments and can be used to perform complex tasks in a small space, for example, administering medicine in human tissues and organs.

[0057] Regarding the dynamic flow field environment in which a magnetically driven microrobot resides and is subject to obstacles, obstacles are physical entities that impede the robot's free movement or interfere with the execution of its mission (e.g., drug delivery). For example, if the magnetically driven microrobot is performing drug delivery within the human body, obstacles may include blood cells, blood vessel walls, thrombi, or other biological structures, without limitation.

[0058] Among them, the state information is used to describe the spatial state of the magnetically driven microrobot and the obstacle, including the real-time distance between the magnetically driven microrobot and the target position, and the positional relationship between the magnetically driven microrobot and the nearest obstacle.

[0059] In one possible implementation, camera sensing and a magnetic control feedback system can be used to acquire the spatial state of the magnetically driven microrobot, thereby obtaining corresponding state information.

[0060] Furthermore, fluid information is used to describe the dynamic flow field environment where the magnetically driven microrobot is located and where obstacles exist. The fluid information can affect the physical environment faced by the magnetically driven microrobot in the dynamic flow field environment, for example, the flow velocity information of the environment where the magnetically driven microrobot is located.

[0061] In one possible implementation, the state information corresponding to the magnetically driven microrobot (for example, real-time position) can be used in combination with a pre-established fluid distribution model of relevant channels in the human body (such as blood vessels, digestive tract, etc.) to map the magnetically driven microrobot to the corresponding flow velocity field, thereby obtaining the corresponding fluid information.

[0062] In one possible implementation, the fluid information may also include flow field distribution characteristics characterized by contrast agents, fluorescent particles or numerical simulation methods to assist in identifying key areas such as the flow direction, backflow area or vortex area where the magnetically driven microrobot is located.

[0063] In addition, micro sensors for collecting fluid information can be integrated into the magnetically driven micro robot, so that fluid information can be obtained from the data returned by the micro sensors.

[0064] In one possible implementation, when a microrobot composed of magnetic material (e.g., a magnetically driven microrobot) is placed in a uniform magnetic field The effect of this will generate a magnetic torque for in, is the magnetization intensity of the object, and the integral is the volume of the magnetized object carried out.

[0065] Furthermore, the magnetically driven microrobot is driven by a uniform rotating field generated by a Helmholtz coil system. The magnetically driven microrobot is subjected to a magnetic torque. The forward speed of the magnetically driven microrobot is 100 times that of the magnetically driven microrobot, and in quasi-static operation, it always tries to align its dipole moment with the frequency of the applied field, thereby achieving forward motion and directional control. and rotation frequency The relationship between them is: r =Kω r ;in, Depends on the geometry and fluid information of the magnetically driven microrobot.

[0066] Step 330 : Based on the state information and the fluid information, the characteristics of the movement of the magnetically driven micro-robot in the dynamic flow field environment with obstacles are captured to obtain the target characteristics.

[0067] First of all, it should be noted that the spatiotemporal characteristics refer to the comprehensive characteristics that describe the changing laws of the magnetically driven microrobot in the spatial dimension, including the position changes of the magnetically driven microrobot in the dynamic flow field environment, the spatial distribution characteristics of the fluid in the magnetically driven microrobot (for example, the changes in flow velocity in different spatial regions, etc.), etc.

[0068] It is understood that state information and fluid information have different characteristics and include different types of features. Different feature extraction methods are suitable for different types of features. Therefore, in order to better extract state information and fluid information, the features of the two should be extracted separately.

[0069] In one possible implementation, MLP (fully connected network) is used to extract features from state information to obtain state features, and the Transforme architecture is used to extract global features from fluid information.

[0070] Of course, in the path planning of the magnetically driven microrobot, the movement of the magnetically driven microrobot depends on its own spatial state and the fluid state of the dynamic flow field environment. Therefore, the features of the two can be fused in order to learn the interaction between the two.

[0071] In one possible implementation, such as Figure 4 As shown, step 330 may further include the following steps:

[0072] Step 331 : Based on the matrix corresponding to the fluid information, the fluid information is converted into a tensor, and the converted fluid information is input into the Transformer module.

[0073] Among them, two layers of the matrix corresponding to the fluid information are used to indicate the key points corresponding to the dynamic flow environment, and each key point is used to indicate the size and direction of the local flow in the dynamic flow field environment; one layer of the matrix corresponding to the fluid information is used to mark the position of the magnetically driven microrobot in the dynamic flow field environment.

[0074] First of all, it should be noted that the number and arrangement of key points are related to the size of the matrix corresponding to the fluid information. For example, the size of the matrix corresponding to the fluid information is 64×64×3. Then, a layer of matrix represents that the dynamic flow environment is divided into 64×64 key points. Each key point corresponds to a vector, which is used to indicate the size and direction of the local flow corresponding to the key point.

[0075] In one possible implementation, a 64×64×3 matrix corresponding to the fluid information is converted into a tensor and then input into the Transformer module. The first two layers of the matrix represent the map corresponding to the dynamic flow environment, divided into 64×64 key points. Each key point has a vector indicating the magnitude and direction of the local flow. The third layer of the matrix marks the position of the magnetically driven microrobot in the discrete dynamic flow field environment.

[0076] In step 333 , a global feature extraction process is performed on the fluid information in the spatial domain through the Transformer module to obtain a global feature.

[0077] First of all, it should be noted that the Transformer module is a deep learning architecture based on the self-attention mechanism, which includes a self-attention mechanism to establish global dependencies in space and identify complex fluid evolution laws.

[0078] Specifically, by utilizing the self-attention mechanism of the Transformer module, we can dynamically capture the long-range dependencies of fluid data in time and space, and thus obtain global features.

[0079] In one possible implementation, the fluid information is processed by three Transformer modules and three fully connected layers to obtain global features.

[0080] Step 335: Perform feature extraction processing on the state information through a multi-layer perceptron to obtain state features.

[0081] Specifically, a multi-layer perceptron (MLP) is used to perform nonlinear mapping and feature abstraction on the state information (such as position, velocity, and posture) of the magnetically driven microrobot to obtain state features.

[0082] In step 337 , the global features and the state features are fused to obtain the target features.

[0083] Among them, fusion processing can be achieved by feature splicing, weighted summation or attention fusion.

[0084] In one possible implementation, the global feature is f c , the state characteristic is f s , for f c and f s Perform splicing processing to obtain the target features.

[0085] Through the above process, by combining the Transformer module and the multi-layer perceptron (MLP), the spatiotemporal characteristics of the magnetically driven microrobot moving in a dynamic flow field environment with obstacles can be well captured, thereby providing more effective information for the subsequent path planning of the magnetically driven microrobot moving in a dynamic flow field environment with obstacles and improving the accuracy of path planning.

[0086] Step 350 , based on the reward function and target features constructed based on the energy consumed by the magnetically driven micro robot during movement in a dynamic flow field environment with obstacles, path planning is performed for the next action of the magnetically driven micro robot to obtain path information.

[0087] Among them, the reward function can be used to quantify the reward signal of the magnetically driven microrobot after performing a certain action. Furthermore, the reward function can affect the path planning of the magnetically driven microrobot. It can be understood that the purpose of path planning can not only be to obtain the maximum reward in a certain action, but also to maximize the cumulative reward that can be achieved after completing the entire path planning (for example, the magnetically driven microrobot reaches the position of the magnetically driven microrobot).

[0088] It can be understood that the construction of the reward function should be consistent with the path planning of the magnetically driven microrobot. Specifically, it can include energy consumption, smoothness, safety, arrival time, etc. Then, a corresponding reward function can be constructed based on the path planning of the magnetically driven microrobot to guide the path planning.

[0089] In one possible implementation, such as Figure 5 As shown, before step 350, the following steps may also be included:

[0090] Step 410: Obtain energy consumption information.

[0091] Among them, the energy consumption information is used to describe the energy consumption of the magnetically driven microrobot moving in a dynamic flow field environment with obstacles.

[0092] It is understandable that since the generation of an external magnetic field requires a large amount of electrical energy (such as a Helmholtz coil), excessive energy consumption may generate excessive heat and affect the safety of surrounding tissues or equipment. Therefore, in order to ensure the safety of the magnetically driven microrobot when moving inside the human body, the energy consumption of the magnetically driven microrobot should be minimized.

[0093] In one possible implementation, before step 410, the following steps may also be included: performing calculation based on the correlation between the propulsion speed generated by the magnetically driven microrobot and the flow velocity in the dynamic flow field environment to obtain the net speed of the magnetically driven microrobot; performing integral calculation based on the net speed of the magnetically driven microrobot to obtain the path passed by the magnetically driven microrobot; determining the energy coefficient of the magnetically driven microrobot; and performing calculation based on the energy coefficient, the propulsion speed generated by the magnetically driven microrobot, and the path passed by the magnetically driven microrobot to obtain energy consumption information.

[0094] First of all, it should be noted that in a dynamic flow field environment with moving obstacles, the environmental forces acting on the magnetically driven microrobot and its dynamics are essentially coupled with each other. This coupling complicates the path planning task, but at the same time, it is also possible to use fluid to plan a path that can save energy.

[0095] Specifically, at time t, the net velocity V of the magnetically driven microrobot is n The resulting propulsion speed V m and flow rate V f Determine, as shown in formula (1):

[0096]

[0097] Wherein, L(t) represents the path passed.

[0098] The energy consumed by the magnetically driven microrobot during the mission can be expressed as a function of its speed, as shown in formula (2):

[0099]

[0100] Among them, m α is the energy coefficient, α∈{1,2,…}. For a magnetically actuated microrobot driven by a Helmholtz coil system, the energy consumption can be quantified by the current consumed in the coil, which is determined by the relationship between the robot's velocity and the applied magnetic field.

[0101] Through the above process, an energy consumption model of the magnetically driven microrobot moving in a dynamic flow field environment was established, which can provide accurate energy consumption information, thereby providing energy consumption constraints for the rewards of subsequent reward functions.

[0102] Step 430: construct the state space and action space.

[0103] The state space includes the spatial state of obstacles in the dynamic flow field environment and the spatial state of the magnetically driven micro-robot; the action space includes the action state of the magnetically driven micro-robot.

[0104] Specifically, for a magnetically driven microrobot navigating in a dynamic environment, the state space It can be expressed as: s={D,V f , P o},in Indicates the position P of the current magnetically driven micro robot r =[x r ,y r ], to the target position P g =[x g ,y g ]’s real-time distance; Indicates that at position P r The local flow velocity at Represents the position of the nearest obstacle relative to the magnetically driven microrobot.

[0105] Furthermore, according to the parameters of the rotating magnetic field, the action space is defined as a={B, f, Θ}, where B, f, Θ represent the intensity, frequency and direction of the generated rotating magnetic field, respectively.

[0106] Step 450 : establishing a smoothness evaluation index based on two consecutive motion vectors of the magnetically driven micro-robot at adjacent time steps in the motion space.

[0107] Among them, the smoothness evaluation index is used to measure whether the movement changes of the magnetically driven microrobot are smooth.

[0108] In order to improve the smoothness of the magnetically driven microrobot path and reduce sudden changes in motion, a smoothness evaluation index S(P) can be introduced. The smoothness evaluation index is composed of two continuous motion vectors P in adjacent time steps.i-1 and P i The smoothness evaluation index is defined by the angle between them. The smoothness evaluation index is shown in formula (3):

[0109]

[0110] Step 470 : construct a corresponding reward function for path planning based on the energy consumption information, the state space, and the smoothness evaluation index.

[0111] Specifically, the reward function is shown in formula (4):

[0112]

[0113] It can be seen from formula (4) that when the magnetically driven micro-robot is less than the safety distance d_s from the nearest obstacle, giving it a penalty of -1000 can enable the magnetically driven micro-robot to learn to avoid obstacles and avoid collisions.

[0114] When the magnetically driven microrobot reaches the vicinity of the target position (the distance is less than d_g), a positive reward of +1500 is given. A high positive reward is given to motivate the magnetically driven microrobot to reach the target position as quickly as possible.

[0115] In other cases, the reward is -E(t)+S(P), where E(t) is the energy consumption and S(P) is the smoothness evaluation index. The smoothness evaluation index can reduce unnecessary turning of the magnetically driven microrobot, making the movement of the magnetically driven microrobot more efficient, and may also reduce mechanical wear or energy consumption. The energy consumption term -E(t) encourages reducing energy use, avoiding heat, and improving safety.

[0116] In combination with the above embodiments, through a multi-dimensional reward and punishment mechanism, high penalties and rewards are used to ensure safety and task priority, and safe, efficient and energy-saving path planning is achieved in a complex dynamic flow field environment. It is especially suitable for scenarios such as medical micro-robots that have strict requirements on reliability and energy efficiency, and can better balance energy efficiency and obstacle avoidance path planning problems.

[0117] It should be noted that the path information can be the action instructions or trajectory points output by the path planning. For example, the path information can include the steering angle, thrust, speed increment and other information of the magnetically driven micro robot. It can also be the position of a series of trajectory points of the magnetically driven micro robot heading to the position of the magnetically driven micro robot. No specific limitation is made here.

[0118] In one possible implementation, such as Figure 6 As shown, step 350 may further include the following steps: inputting the target feature into the SAC module, performing path planning for the next action of the target based on the target feature and the reward function, and obtaining path information; wherein the SAC module includes a policy network.

[0119] First of all, it should be noted that the SAC module is an actor-critic method that combines policy entropy to encourage exploration. The goal is to maximize the expected reward and policy entropy. The formula of the SAC module is shown in formula (5):

[0120]

[0121] Among them, R(s t , a t ) is the reward of the current state-action pair, γ∈(0,1] is the discount factor, and α is the temperature parameter that balances the reward and the policy entropy H.

[0122] Furthermore, step 350 may include:

[0123] Step 351: Input the target features into the strategy network to determine the current state of the magnetically driven microrobot.

[0124] First of all, it should be noted that the target characteristics integrate the global characteristics of the dynamic flow field environment and the state characteristics of the magnetically driven microrobot. Therefore, the current state of the magnetically driven microrobot can be determined based on the target characteristics. The current state can describe the comprehensive environmental state of the magnetically driven microrobot, including the current posture of the magnetically driven microrobot, its position in the dynamic flow field environment, etc., which are not limited here.

[0125] Step 353: output the action probability distribution of the magnetically driven micro robot based on the current state, and sample the action probability distribution to obtain the next action of the magnetically driven micro robot in the process of moving to the target position.

[0126] First of all, it should be noted that the policy network has the characteristics of Gaussian distribution parameterization, so the policy network can observe the current state s t , from Ω η (a t |s t ) Select an action a t , where a t This is the next action of the magnetically driven microrobot when it moves to the target position. η (a t |s t ) is the action policy distribution of the policy network.

[0127] It is further explained that the SAC module also includes a target network and a dual critic network. In a possible implementation, Figure 7 As shown, step 353 may further include the following steps:

[0128] Step 510 : Process the current state and next action of the magnetically driven micro robot according to the reward function to generate corresponding transfer samples.

[0129] The transferred samples include the current state, next action, environmental reward and next state of the magnetically driven microrobot.

[0130] Regarding the environmental reward, the reward function constructed in the above embodiment can be used to obtain the reward based on the current state. Regarding the next action, it is determined by the current state and the next action of the magnetically driven micro robot.

[0131] In one possible implementation, from Ω η (a t |s t ) Select an action a t , and transfer (s t , a t , r t , s t+1 , d t ) is stored in the experience replay buffer D, where r t For environmental rewards, S t+1 is the next state, d t It is a termination flag used to indicate whether the magnetically driven microrobot has reached the target position or collided with an obstacle.

[0132] In step 530 , the transfer sample is input into the dual critic network, and the value is estimated according to the reward function to obtain the action-value function.

[0133] Among them, the action-value function is used to describe the expected long-term reward that the magnetically driven microrobot can obtain by performing the next action in the current state.

[0134] First, it should be noted that since action-value function estimation in reinforcement learning may have overestimation bias, by using two critic networks and taking the minimum action-value function of the two when training the policy network, this bias can be effectively suppressed, thereby enhancing the stability and reliability of the policy.

[0135] It is additionally noted that after step 530, the following steps may also be included:

[0136] Based on the action-value function, the policy network is optimized by maximizing the entropy enhancement strategy.

[0137] It should be noted that the use of two critic networks can reduce the over-estimation bias during policy learning. Each critic network independently estimates an action value function, and the smaller of the two is used to update the policy network.

[0138] Specifically, the policy network (Actor-network) is updated by maximizing the entropy enhancement policy objective, as shown in formula (7):

[0139]

[0140] Among them, J π (θ) is the optimization goal of the policy network, Q φ (s t , a t ) is the action value function, Ω η (a t |s t ) is the action probability distribution of the current policy network, Indicates the current state s t expected value.

[0141] Through the above process, the policy goal is enhanced by maximizing entropy, which encourages policy exploration, encourages diversified actions, finds better paths in complex flow fields, and avoids premature convergence to local optimality.

[0142] In step 550 , the action-value function and the transfer sample are input into the target network, and the critic network is optimized according to the Bellman residual minimization method.

[0143] First, when using a critic network, training directly with a real-time updated network can lead to instability or even training divergence. Therefore, a target network is introduced as a stable reference. The target network is not updated every time, but rather its weights are gradually updated using a delayed parameter method (soft update) to maintain stability.

[0144] Specifically, the target network is used to calculate a stable target value. The target network can stabilize value estimation and avoid training divergence by delaying parameter updates. In addition, the minimum value of the two critic networks is taken as the action-value function to reduce over-estimation error. By coordinating the dual critic network with the target network, the overestimation problem of the action-value function can be suppressed.

[0145] The critic network performs training optimization by comparing the gap (Bellman residual) between its own predicted action-value function and the target Q-value estimated by the target network based on the next state. The goal is to make the action-value function output by the critic network closer to the "true" long-term reward estimate.

[0146] Specifically, the critic network is trained by minimizing the soft Bellman residual, as shown in formula (8):

[0147]

[0148] Among them, J Q (φ) represents the loss function of the critic network, φ is the parameter of the critic network, Expressing expectations, (s t , a t , r t , s t+1 , d t ) is the transfer sample sampled from the experience replay buffer D, Q φ (s t , a t ) is the current critic network's response to the current state s t and the next action a t The action value function obtained by estimating the value of is the target Q value calculated by the target network.

[0149] Among them, the target Q value can be based on Calculation, r t To take action a at time t t The reward obtained later; γ is the discount factor, which usually ranges from (0, 1] and is used to measure the importance of future rewards. The closer γ is to 1, the more importance is attached to future rewards; d t It is a termination flag used to indicate whether the magnetically driven micro robot has reached the termination state at time t (if it reaches the termination state, d t =1, otherwise d t =0); V(s t+1 ) is the target network's response to the next state s t+1 estimated value.

[0150] Through the above process, the target network can update and optimize the critic network, ensuring that the subsequently generated value estimation function is closer to the actual situation, and thus use the more accurate value estimation function to update the policy network, making the next action it outputs more precise, thereby improving the safety and accuracy of path planning.

[0151] Combined with the above embodiments, SAC realizes intelligent path planning of magnetically driven microrobots in complex dynamic environments through a ternary collaborative mechanism of dual critic network value evaluation, target network stable training, and policy network optimization of action distribution. It can maximize the entropy of the strategy while maximizing the expected return, thereby encouraging exploration and avoiding falling into local optimality.

[0152] Specifically, path planning needs to take into account factors such as dynamic flow fields, obstacle avoidance, energy efficiency, and path smoothness. The entropy maximization mechanism of SAC can enhance the exploration ability of magnetically driven microrobots in complex environments, help them find better paths, and ensure that they can achieve precise obstacle avoidance and energy consumption control in high-risk environments. In dynamic flow fields, SAC can also adjust strategies in real time to cope with environmental changes, in order to cope with complex scenarios such as sudden changes in flow rate and obstacle drift.

[0153] Step 370 : Control the magnetically driven micro-robot to move in a dynamic flow field environment with obstacles according to the path information, so that the magnetically driven micro-robot approaches or reaches the target position.

[0154] The target position is the position that the magnetically driven microrobot needs to reach to perform a task. For example, if the magnetically driven microrobot performs a drug delivery task, the target position may be the lesion.

[0155] It should be noted that path planning is a dynamic process. Every time the magnetically driven microrobot performs an action, the position of the magnetically driven microrobot should be close to the target position. Therefore, based on continuous path planning, the magnetically driven microrobot performs a series of actions and finally reaches the target position.

[0156] Then, based on the next action indicated by the path information, a corresponding control signal can be generated based on the magnetically driven microrobot, so that the magnetically driven microrobot can perform the next action according to the control signal.

[0157] Through the above process, the reward function and target characteristics are constructed based on the energy consumed by the magnetically driven microrobot during its movement in a dynamic flow field environment with obstacles, and the path planning of the next action of the magnetically driven microrobot is carried out. The fluid velocity field, obstacle distribution and energy consumption are taken into account. The direction and intensity of the magnetic field are adjusted in real time through the magnetic field control system, so that the magnetically driven microrobot can avoid obstacles along the path, overcome fluid resistance, reduce energy consumption, and move to the target position accurately, ensuring the completion and safety of the task.

[0158] Figures 8a to 8c This is a schematic diagram of a specific implementation of a path planning method in an application scenario. This application scenario simulates a magnetically driven microrobot delivering drugs inside the human body, where glycerol is used to simulate the dynamic flow field environment inside the human body.

[0159] Figure 8a A flowchart of the application scenario is shown. Figure 8b The structural block diagram of the path planning model in this application scenario is shown.

[0160] Now combined Figure 8a and Figure 8b The following describes the application scenario:

[0161] First, the environment is initialized, that is, the starting point, target position, dynamic flow field environment and obstacles of the magnetically driven microrobot are set.

[0162] The camera obtains the status information of the magnetically driven microrobot and the fluid information, and inputs the status information and the fluid information into the path planning model.

[0163] Specifically, the 64×64×3 matrix corresponding to the fluid information is converted into a tensor and then input into the Transformer module. Three Transformer modules and three fully connected layers perform global feature extraction on the fluid information in the spatial domain to obtain global features. A multilayer perceptron extracts features from the state information to obtain state features. Finally, the state features and global features are concatenated to obtain the target features, which are then input into the SAC module.

[0164] The SAC module constructs a reward function based on the energy consumed by the magnetically driven microrobot during its movement in a dynamic flow field environment with obstacles, as well as target features, to plan the path for the next action of the magnetically driven microrobot and obtain path information.

[0165] Specifically, the target feature is input into the strategy network of the SAC module to determine the current state of the magnetically driven microrobot. The strategy network outputs the action probability distribution of the magnetically driven microrobot based on the current state, and samples the action probability distribution to obtain the next action of the magnetically driven microrobot in the process of moving to the target position, thereby obtaining the path information, such as Figure 8c As shown, Figure 8c The figure shows a specific implementation diagram of path planning in this application scenario.

[0166] According to the path information, the three-dimensional Helmholtz coil drives the magnetically driven microrobot to move in a dynamic flow field environment to approach or reach the lesion (target position). If the magnetically driven microrobot has not reached the lesion after completing the next action, path planning will continue based on the current state of the magnetically driven microrobot until the magnetically driven microrobot reaches the lesion.

[0167] In this application scenario, by using Transformer to extract the global spatiotemporal characteristics of the flow field and combining it with the SAC algorithm for adaptive decision-making, the functions of balancing energy efficiency, path continuity and dynamic obstacle avoidance are achieved, and a path that is both energy-saving and effectively obstacle-avoiding is generated in the dynamic flow field. The magnetically driven micro-robot is guided to reach the target position stably, so that the magnetically driven micro-robot can avoid obstacles and reach the target position with lower energy consumption and more stability in the dynamic flow field, solving the difficulties faced by traditional methods in dynamic flow field path planning.

[0168] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0169] The following are embodiments of the device of the present application, which can be used to execute the path planning method involved in the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the path planning method involved in the present application.

[0170] See also Figure 9 In an embodiment of the present application, a path planning device 900 is provided, including but not limited to: an information acquisition module 910 , a feature processing module 930 , a path planning module 930 and a motion control module 950 .

[0171] The information acquisition module 910 is used to acquire state information and fluid information; the state information is used to describe the spatial state of the magnetically driven micro-robot; the fluid information is used to describe the dynamic flow field environment where the magnetically driven micro-robot is located and has obstacles;

[0172] The feature processing module 930 is used to capture the spatiotemporal characteristics of the movement of the magnetically driven micro-robot in a dynamic flow field environment with obstacles based on the state information and fluid information to obtain the target characteristics;

[0173] A path planning module 950 is configured to plan a path for the next action of the magnetically driven microrobot based on a reward function and target characteristics constructed based on the energy consumed by the magnetically driven microrobot during movement in a dynamic flow field environment with obstacles, thereby obtaining path information.

[0174] The motion control module 950 is used to control the magnetically driven micro robot to move to a target position in a dynamic flow field environment with obstacles according to the path information.

[0175] It should be noted that the path planning device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when performing path planning. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the path planning device will be divided into different functional modules to complete all or part of the functions described above.

[0176] In addition, the path planning device and the path planning method provided in the above embodiments belong to the same concept, and the specific manner in which each module performs operations has been described in detail in the method embodiments and will not be repeated here.

[0177] See also Figure 10 In an embodiment of the present application, an electronic device 4000 is provided. The electronic device 4000 may include: a desktop computer, a laptop computer, a server, etc.

[0178] exist Figure 10 In the embodiment, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.

[0179] Data exchange between the processor 4001 and the memory 4003 can be achieved via at least one communication bus 4002. The communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0180] Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0181] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0182] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store a computer program in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to these.

[0183] The memory 4003 stores a computer program, and the processor 4001 can read the computer program stored in the memory 4003 through the communication bus 4002 .

[0184] The computer program is executed by one or more processors 4001 to implement the path planning method in the above embodiments.

[0185] In addition, an embodiment of the present application provides a storage medium on which a computer program is stored. The computer program is executed by one or more processors to implement the path planning method described above.

[0186] A computer program product is provided in an embodiment of the present application, including a computer program, which is executed by one or more processors to implement the path planning method described above.

[0187] Compared with related technologies, the reward function and target characteristics constructed based on the energy consumed by the magnetically driven microrobot during its movement in a dynamic flow field environment with obstacles are used to plan the path of the next action of the magnetically driven microrobot, taking into account the fluid velocity field, obstacle distribution and energy consumption. The magnetic field direction and intensity are adjusted in real time through the magnetic field control system, so that the magnetically driven microrobot can avoid obstacles along the path, overcome fluid resistance, reduce energy consumption, and move to the target position accurately, ensuring the completion and safety of the task.

[0188] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A path planning method for a magnetically driven microrobot, characterized in that: include: Obtain status information and fluid information; The state information is used to describe the spatial state of the magnetically driven micro-robot and the obstacle; The fluid information is used to describe the dynamic flow field environment in which the magnetically driven microrobot is located; Based on the state information and the fluid information, capturing the movement characteristics of the magnetically driven microrobot in the dynamic flow field environment with obstacles to obtain target characteristics; performing path planning for the next action of the magnetically driven microrobot based on a reward function constructed based on the energy consumed by the magnetically driven microrobot during movement in the dynamic flow field environment with obstacles, and the target feature, to obtain path information; The magnetically driven microrobot is controlled to move in the dynamic flow field environment where obstacles exist according to the path information, so that the magnetically driven microrobot approaches or reaches the target position.

2. The method according to claim 1, wherein The capturing, based on the state information and the fluid information, the movement characteristics of the magnetically driven microrobot in the dynamic flow field environment with obstacles to obtain target characteristics includes: Based on the matrix corresponding to the fluid information, the fluid information is converted into a tensor, and the converted fluid information is input into a Transformer module; wherein two layers of the matrix corresponding to the fluid information are used to indicate key points corresponding to the dynamic flow environment, and each key point is used to indicate the size and direction of the local flow in the dynamic flow field environment; one layer of the matrix corresponding to the fluid information is used to mark the position of the magnetically driven microrobot in the dynamic flow field environment; Performing global feature extraction processing on the fluid information in the spatial domain through a Transformer module to obtain global features; Performing feature extraction processing on the state information through a multi-layer perceptron to obtain state features; The global feature and the state feature are fused to obtain the target feature.

3. The method according to claim 1, wherein The reward function constructed based on the energy consumed by the magnetically driven micro-robot during movement in the dynamic flow field environment with obstacles and the target feature is used to perform path planning for the next action of the magnetically driven micro-robot to obtain path information, including: Inputting the target feature into the SAC module, and performing path planning for the next action of the target based on the target feature and the reward function to obtain the path information; wherein the SAC module includes a policy network; Inputting the target feature into the SAC module, and performing path planning for the next action of the target based on the target feature and the reward function to obtain the path information includes: Input the target features into the policy network, An action probability distribution is output based on the current state of the magnetically driven micro robot, and sampling is performed in the action probability distribution to obtain the next action of the magnetically driven micro robot in the process of moving to the target position.

4. The method according to claim 3, wherein The SAC module also includes a target network and a dual critic network; After outputting the action probability distribution of the magnetically driven micro-robot based on the current state and sampling in the action probability distribution to obtain the next action of the magnetically driven micro-robot in the process of moving to the target position, the method further includes: Processing the current state and the next action of the magnetically driven microrobot according to the reward function to generate a corresponding transfer sample; the transfer sample includes the current state of the magnetically driven microrobot, the next action, the environmental reward, and the next state; Inputting the transfer sample into the dual critic network, performing value estimation based on the reward function, and obtaining an action-value function; the action-value function is used to describe the expected long-term reward that can be obtained by the magnetically driven micro robot performing the next action in the current state; The action-value function and the transfer sample are input into the target network, and the dual-critic network is optimized according to the Bellman residual minimization method.

5. The method according to claim 4, wherein After inputting the transfer sample into the critic network for value estimation and obtaining the action-value function, the method further includes: Based on the action-value function, the policy network is optimized by maximizing entropy enhancement strategy.

6. The method according to any one of claims 1 to 5, characterized in that Before obtaining the path information, the method further includes: planning a path for the next action of the magnetically driven micro-robot based on the reward function constructed based on the energy consumed by the magnetically driven micro-robot during movement in the dynamic flow field environment with obstacles and the target feature. Obtaining energy consumption information; the energy consumption information is used to describe the energy consumption of the magnetically driven micro-robot moving in a dynamic flow field environment with obstacles; Constructing a state space and an action space; the state space includes the spatial state of obstacles in the dynamic flow field environment and the spatial state of the magnetically driven micro-robot; the action space includes the action state of the magnetically driven micro-robot; Establishing a smoothness evaluation index based on two continuous motion vectors of the magnetically driven microrobot at adjacent time steps in the motion space; the smoothness evaluation index is used to measure whether the motion change of the magnetically driven microrobot is smooth; Based on the energy consumption information, the state space and the smoothness evaluation index, a corresponding reward function is constructed for the path planning.

7. The method according to claim 6, wherein Before obtaining the energy consumption information, the method further includes: Calculating the net speed of the magnetically driven microrobot based on the correlation between the propulsion speed generated by the magnetically driven microrobot and the flow speed of the dynamic flow field environment; Performing an integral calculation based on the net speed of the magnetically driven micro-robot to obtain a path traversed by the magnetically driven micro-robot; determining an energy coefficient of the magnetically driven microrobot; The energy consumption information is obtained by performing calculation based on the energy coefficient, the propulsion speed generated by the magnetically driven microrobot, and the path passed by the magnetically driven microrobot.

8. A path planning device for a magnetically driven microrobot, characterized in that: include: An information acquisition module is used to obtain status information and fluid information; The state information is used to describe the spatial state of the magnetically driven micro-robot and the obstacle; The fluid information is used to describe the dynamic flow field environment where the magnetically driven microrobot is located; a feature processing module, configured to capture, based on the state information and the fluid information, features of the magnetically driven microrobot moving in the dynamic flow field environment with obstacles, and obtain target features; a path planning module, configured to plan a path for the next action of the magnetically driven microrobot based on a reward function constructed based on the energy consumed by the magnetically driven microrobot during movement in the dynamic flow field environment with obstacles and the target features, and obtain path information; A motion control module is used to control the magnetically driven micro robot to move to a target position in the dynamic flow field environment where obstacles exist according to the path information.

9. An electronic device comprising at least one processor and at least one memory, wherein: A computer program is stored in the memory, wherein when the computer program is executed by the processor, the path planning method according to any one of claims 1 to 7 is implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by one or more processors, the path planning method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Magnetic drive water surface robot position control method and system based on fuzzy adaptive LQR

    CN121704169A