A Reinforcement Learning-Based AGV Transportation Path Planning Method and System

By constructing a three-dimensional dynamic electromagnetic field model and a safety barrier reinforcement learning decision framework, the problem of balancing transportation efficiency and electromagnetic safety of AGVs in high-voltage electrical detection environments was solved, and safe and efficient path planning was achieved.

CN122505262APending Publication Date: 2026-08-04GANZHOU JIANGYUAN TECHNOLOGY DEVELOPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GANZHOU JIANGYUAN TECHNOLOGY DEVELOPMENT CO LTD
Filing Date
2026-05-08
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing AGV transportation path planning methods are difficult to simultaneously balance transportation efficiency and electromagnetic safety constraints in high-voltage electrical detection environments.

Method used

Electromagnetic data is collected collaboratively by a fixed electromagnetic field sensor array and an automated guided vehicle (AGV), a three-dimensional dynamic electromagnetic field model is constructed, and an electromagnetic situation heatmap is generated. Combined with the state data of the AAV, a composite state space is constructed, and a safety barrier reinforcement learning decision framework is adopted, consisting of an upper-layer reinforcement learning policy network and a lower-layer safety barrier corrector, to verify and correct the path planning in real time.

Benefits of technology

It achieves high-precision electromagnetic environment perception, ensuring safe and efficient transportation of AGVs in complex environments, and improving the safety and intelligence level of autonomous transportation operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122505262A_ABST
    Figure CN122505262A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for AGV (Automated Guided Vehicle) transportation path planning based on reinforcement learning, belonging to the field of reinforcement learning technology. The method includes: loading electromagnetic situational awareness data from a detection workshop, constructing a three-dimensional dynamic electromagnetic field spatiotemporal distribution model and generating an electromagnetic situational heatmap; loading geometric, kinematic, and task state data of the automated guided vehicle (AGV) to construct a composite state space; constructing a safety barrier reinforcement learning decision framework including an upper-layer reinforcement learning policy network and a lower-layer safety barrier corrector, where the upper-layer network outputs a preliminary desired action, and the lower-layer corrector performs safety verification; when the preliminary desired action passes the verification, it is output as an actual control command; when it fails, a safety correction action is generated as the actual control command, and the verification failure event is fed back to the upper-layer network as a penalty for learning. This invention achieves a balance between transportation efficiency and electromagnetic safety in automated guided vehicle (AGV) transportation path planning under high-voltage electrical detection environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reinforcement learning technology, and specifically to an AGV transportation path planning method and system based on reinforcement learning. Background Technology

[0002] In high-voltage electrical performance testing scenarios, the transportation route planning of AGVs (Automated Guided Vehicles) must simultaneously consider transportation efficiency and electromagnetic safety constraints.

[0003] Existing technologies, such as CN113128770B, disclose a real-time optimization method for material distribution in uncertain workshop environments based on DQN. This method optimizes the distribution path through dynamic time windows and path resistance coefficients, but it does not consider the impact of dynamic electromagnetic fields generated by high-voltage equipment on the AGV control system or safety clearance requirements. Another example is CN121073708B, which discloses a method for generating safety protection paths for power personnel based on multi-objective optimization. This method utilizes drones to monitor the electromagnetic distribution of substations and generate safety protection paths. However, this method focuses on personnel movement and does not involve the reinforcement learning decision-making framework for AGV transportation tasks, nor does it provide real-time hard constraints on electromagnetic safety red lines during AGV operation.

[0004] Therefore, existing AGV transportation path planning methods are difficult to implement in high-voltage electrical detection environments, where both efficiency and electromagnetic safety are considered. Summary of the Invention

[0005] This invention addresses the technical problem in existing technologies where automated guided vehicles (AGVs) path planning methods struggle to simultaneously balance transportation efficiency and electromagnetic safety constraints in complex detection environments containing high-voltage electrical equipment. It provides an AGV path planning method and system based on reinforcement learning.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: In a first aspect, the present invention provides an AGV transportation path planning method based on reinforcement learning, comprising: The electromagnetic situational awareness data of the testing workshop is loaded, wherein the electromagnetic situational awareness data includes first field strength data collected by a fixed electromagnetic field sensor array and second field strength data collected by an automated guided vehicle. Based on the electromagnetic situational awareness data, a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the testing workshop is constructed to obtain an electromagnetic situational heat map superimposed on a two-dimensional grid map, wherein each grid point in the electromagnetic situational heat map corresponds to a real-time field strength value. Load the geometric state data, kinematic state data, and mission state data of the automated guided vehicle, and combine them with the electromagnetic situation heat map to construct a composite state space; Based on the composite state space, a safety barrier reinforcement learning decision framework is constructed, which includes an upper-layer reinforcement learning policy network and a lower-layer safety barrier modifier. The upper-layer reinforcement learning policy network outputs a preliminary expected action, and the lower-layer safety barrier modifier performs real-time safety verification on the preliminary expected action based on an electromagnetic safety threshold. When the underlying security barrier modifier determines that the preliminary expected action passes the security check, it outputs the preliminary expected action as the actual control command. When the underlying security barrier modifier determines that the initial expected action has failed the security check, it generates a minimum intervention security correction action, uses the security correction action as the actual control instruction, and feeds back the event of this security check failure as a penalty to the upper-layer reinforcement learning policy network for learning. The actual control commands are executed to complete the path planning for the automated guided vehicle.

[0007] Secondly, this invention provides an AGV transportation path planning system based on reinforcement learning, comprising: The data loading module is used to load electromagnetic situational awareness data from the testing workshop. The electromagnetic situational awareness data includes first field strength data collected by a fixed electromagnetic field sensor array and second field strength data collected by an automated guided vehicle. The electromagnetic situation modeling module is used to construct a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the testing workshop based on the electromagnetic situation perception data, and obtain an electromagnetic situation heat map superimposed on a two-dimensional grid map, wherein each grid point in the electromagnetic situation heat map corresponds to a real-time field strength value. The composite state space construction module is used to load the geometric state data, kinematic state data and mission state data of the automated guided vehicle, and combine them with the electromagnetic situation heat map to construct a composite state space; The decision framework construction module is used to construct a safety barrier reinforcement learning decision framework based on the composite state space, which includes an upper-layer reinforcement learning policy network and a lower-layer safety barrier modifier. The upper-layer reinforcement learning policy network outputs a preliminary expected action, and the lower-layer safety barrier modifier performs real-time safety verification on the preliminary expected action based on an electromagnetic safety threshold. The first instruction output module is used to output the preliminary expected action as an actual control instruction when the underlying security barrier modifier determines that the preliminary expected action has passed the security check. The second instruction output and feedback module is used to generate a minimum intervention safety correction action when the underlying security barrier corrector determines that the initial expected action has failed the security check. The safety correction action is used as the actual control instruction, and the event of failing the security check is fed back to the upper-layer reinforcement learning policy network as a penalty for learning. The path planning execution module is used to execute the actual control commands and complete the path planning of the automated guided vehicle.

[0008] The beneficial effects of this invention are: Compared to existing technologies, this invention firstly collects electromagnetic data collaboratively with fixed sensors and automated guided vehicles (AGVs) to construct a three-dimensional dynamic electromagnetic field model and generate an electromagnetic situation heatmap, achieving high-precision perception of the electromagnetic environment in the inspection workshop. Secondly, this invention constructs a composite state space that integrates vehicle status and real-time electromagnetic situation, providing multi-dimensional safety input for reinforcement learning decision-making. Thirdly, this invention proposes a decision framework comprising an upper-level reinforcement learning strategy network and a lower-level safety barrier corrector. The upper-level strategy network is responsible for improving transportation efficiency, while the lower-level safety barrier corrector performs hard safety checks on preliminary expected actions based on electromagnetic safety thresholds, ensuring that the AGV always operates within the electromagnetic safety boundary. Simultaneously, this invention feeds back events that fail safety checks as penalties to the upper-level strategy network for learning, prompting it to proactively avoid dangers. Finally, this invention solves the problem of AGV path planning in high-voltage electrical inspection environments, which struggles to balance transportation efficiency and electromagnetic safety, improving the safety and intelligence level of autonomous transportation operations in complex environments. Attached Figure Description

[0009] Figure 1 A flowchart illustrating the reinforcement learning-based AGV transportation path planning method provided by this invention; Figure 2 This is a schematic diagram of the structure of the AGV transportation path planning system based on reinforcement learning provided by the present invention.

[0010] In the attached diagram, the components represented by each number are as follows: The module includes a data loading module 11, an electromagnetic situation modeling module 12, a composite state space construction module 13, a decision framework construction module 14, a first instruction output module 15, a second instruction output and feedback module 16, and a path planning execution module 17. Detailed Implementation

[0011] Example 1, as Figure 1 As shown, this embodiment of the invention provides an AGV transportation path planning method based on reinforcement learning, including: S10: Load the electromagnetic situational awareness data of the testing workshop, wherein the electromagnetic situational awareness data includes the first field strength data collected by the fixed electromagnetic field sensor array and the second field strength data collected by the automated guided vehicle. First, the electromagnetic situational awareness data from the testing workshop is loaded. The testing workshop refers to a specific work area where high-voltage electrical equipment performance testing is conducted. This area contains operating high-voltage charged bodies, creating a complex dynamic electromagnetic field environment. The electromagnetic situational awareness data is a set of real-time field strength information collected and aggregated through a fixed array of sensors and a moving automated guided vehicle. This data is used to characterize the electromagnetic field distribution intensity and changing trends at different spatial locations within the testing workshop at different times.

[0012] Specifically, the electromagnetic situational awareness data from the testing workshop is loaded, including: The fixed electromagnetic field sensor array is loaded with the first field strength data collected at a preset sampling frequency; The second field strength data collected during the operation of the automated guided vehicle is loaded by the mobile field strength monitoring terminal mounted on the automated guided vehicle. The first field strength data and the second field strength data are aggregated to the central processor via a wireless communication network and added to the electromagnetic situational awareness data.

[0013] First, the first electromagnetic field strength data is collected by a fixed electromagnetic field sensor array at a preset sampling frequency. This fixed electromagnetic field sensor array is pre-deployed at key locations within the testing workshop, such as around high-voltage electrical testing equipment and at passageway intersections. Each sensor in this fixed electromagnetic field sensor array operates continuously at the set sampling frequency, sensing the electromagnetic field strength at its location in real time and generating a corresponding electrical signal. The preset sampling frequency is a pre-defined time interval parameter, representing the number of times the sensor collects electromagnetic field strength data per unit time. The value of this preset sampling frequency is set according to the intensity of electromagnetic field changes within the testing workshop and the system's real-time data requirements; for example, it can be set to collect data once per second.

[0014] After analog-to-digital conversion, the electrical signal forms the first field strength data. This first field strength data is used to reflect the baseline information of the change of electromagnetic field strength over time at fixed spatial points in the detection workshop, providing a stable reference point for constructing a basic model of the static electromagnetic field distribution in the workshop.

[0015] Secondly, a mobile field strength monitoring terminal mounted on the automated guided vehicle (AGV) collects a second field strength data during the AGV's operation. Specifically, the mobile field strength monitoring terminal is installed on the AGV and traverses different areas within the workshop as the AGV moves. During the AGV's transport mission, this monitoring terminal continuously collects the electromagnetic field strength at the locations it passes through, generating the second field strength data. This second field strength data effectively captures areas outside the coverage of fixed sensor arrays and local distortions in the electromagnetic field caused by changes in equipment operating conditions or object movement. It serves as an important supplement to fixed-point measurement data, improving the spatial resolution and dynamic response capability of electromagnetic situational awareness.

[0016] Subsequently, the first and second electromagnetic field strength data are converged to the central processor via a wireless communication network and added to the electromagnetic situational awareness (ESA) data. Both the fixed sensor array and the mobile monitoring terminal are equipped with wireless communication modules. The collected first and second electromagnetic field strength data are transmitted to the central processor via a wireless communication network deployed within the testing workshop. This central processor is a computing device with data receiving, parsing, fusion, and storage functions. It is responsible for receiving, parsing, and fusing data streams from different sources, integrating and storing data marked with timestamps and location information, and ultimately forming complete and real-time updated ESA data for subsequent modeling and analysis steps.

[0017] Specifically, the final electromagnetic situational awareness data is a comprehensive, high-precision set of dynamic electromagnetic field information. It includes first field strength data, collected by a fixed electromagnetic field sensor array providing a global, continuous reference field strength distribution, and second field strength data, collected by a mobile field strength monitoring terminal mounted on an automated guided vehicle (AGV) to compensate for blind spots at fixed points and reflect local dynamic changes. This electromagnetic situational awareness data is used to comprehensively and in real-time characterize the global electromagnetic field distribution within the testing workshop, from macroscopic references to microscopic dynamics.

[0018] S20: Based on the electromagnetic situational awareness data, construct a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the detection workshop, and obtain an electromagnetic situational heat map superimposed on a two-dimensional grid map, wherein each grid point in the electromagnetic situational heat map corresponds to a real-time field strength value. Secondly, based on the aforementioned electromagnetic situational awareness data, a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the testing workshop is constructed. This three-dimensional dynamic electromagnetic field spatiotemporal distribution model is a digital field distribution expression with three-dimensional spatial coordinates and time coordinates as independent variables and electromagnetic field intensity as the dependent variable. It is used to accurately describe the electromagnetic field intensity and its dynamic evolution at any spatial location within the testing workshop at different times.

[0019] Simultaneously, an electromagnetic situation heatmap is obtained overlaid on a two-dimensional grid map. The two-dimensional grid map refers to a digital map where the physical ground area of ​​the testing workshop is divided into uniform grids according to a set resolution, and each grid cell is assigned positional coordinates. The electromagnetic situation heatmap is a graphical representation generated by slicing a three-dimensional dynamic electromagnetic field spatiotemporal distribution model according to the activity height of the automated guided vehicle (AGV), and then mapping the field strength data on the slices to each corresponding grid point on the two-dimensional grid map using color depth or numerical values. Each grid point in this electromagnetic situation heatmap corresponds to a real-time field strength value, providing intuitive and quantified electromagnetic safety constraint information down to each spatial location point for AGV path planning.

[0020] Specifically, based on the electromagnetic situational awareness data, a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the detection workshop is constructed to obtain an electromagnetic situational heat map superimposed on a two-dimensional grid map, including: Based on the principle of finite element analysis, a basic model of the static electromagnetic field distribution of the high-voltage electrical testing equipment in the testing workshop is established. Based on the electromagnetic situational awareness data, the static electromagnetic field distribution model is dynamically corrected by a data fusion algorithm to obtain the three-dimensional dynamic electromagnetic field spatiotemporal distribution model. Based on the height data of the automated guided vehicle and the two-dimensional grid map, the activity space of the automated guided vehicle is obtained; Based on the activity space of the automated guided vehicle, the spatiotemporal distribution model of the target electromagnetic field is extracted from the three-dimensional dynamic electromagnetic field spatiotemporal distribution model; The spatiotemporal distribution model of the target electromagnetic field is mapped onto a two-dimensional grid map of the testing workshop to generate the electromagnetic situation heat map, wherein each grid point in the electromagnetic situation heat map corresponds to a real-time field strength value.

[0021] First, based on the principle of finite element analysis, a basic model of the static electromagnetic field distribution of high-voltage electrical testing equipment in the testing workshop is established. The finite element analysis principle is a numerical analysis method that discretizes a continuous solution domain into a finite number of interconnected micro-elements, constructs an approximate function on each element, and then obtains the physical field distribution over the entire domain by solving the overall system of equations.

[0022] Specifically, the solution domain is first discretized into a finite number of tiny units based on the geometric dimensions, material properties, layout, and rated voltage and current parameters of the high-voltage electrical testing equipment within the testing workshop. Then, the electromagnetic field equations for each tiny unit are constructed using Maxwell's equations, and boundary conditions consistent with actual operating conditions are applied for numerical solution. Maxwell's equations are a set of partial differential equations describing the relationship between electric and magnetic fields and charge and current densities, forming the foundation of classical electromagnetic theory. The boundary conditions consistent with actual operating conditions are specific constraints set for the boundaries of the solution domain, such as a boundary condition where the field strength is zero at infinity, or a boundary condition where the electric field direction is perpendicular to the surface of a conductor. This process calculates the theoretical electromagnetic field strength distribution at various spatial points within the testing workshop under ideal steady-state conditions, thus constructing a static electromagnetic field distribution model reflecting the inherent electromagnetic characteristics of the equipment. This static electromagnetic field distribution model provides a reference initial model for subsequent dynamic corrections.

[0023] Secondly, based on electromagnetic situational awareness data, a data fusion algorithm is used to dynamically correct the static electromagnetic field distribution model, resulting in a three-dimensional dynamic spatiotemporal electromagnetic field distribution model. Specifically, the electromagnetic situational awareness data, collected in real time by the central processor—namely, the first field strength data acquired by the fixed sensor array and the second field strength data acquired by the mobile automated guided vehicle—is used as the measured data input. Simultaneously, data fusion algorithms such as Kalman filtering, particle filtering, or neural networks are employed to fuse the measured discrete-point field strength data with the static electromagnetic field distribution model.

[0024] Specifically, this fusion process uses measured data to dynamically correct the boundary conditions or internal parameters of the basic model, thereby continuously updating the model and eliminating the deviation between theoretical calculations and actual conditions. This allows the model to reflect in real time the dynamic evolution of the electromagnetic field caused by factors such as changes in equipment load, movement of personnel or objects, and ultimately generates a three-dimensional spatiotemporal distribution model of the electromagnetic field that changes dynamically over time. This three-dimensional spatiotemporal distribution model of the electromagnetic field can reflect the dynamic changes of the electromagnetic field within the detection workshop in real time.

[0025] Furthermore, based on the height data of the automated guided vehicle (AGV) and a two-dimensional grid map, the AGV's operating space is obtained. Since the AGV also occupies a certain physical space when it is running in the testing workshop, and the field strength experienced by the vehicle body at different heights in the electromagnetic field environment is also different, it is necessary to provide an electromagnetic safety assessment area that matches the actual contour of the vehicle body for path planning.

[0026] The height data of the automated guided vehicle (AGV) refers to the vertical distance between the top of the vehicle and the ground. The two-dimensional grid map defines the feasible area of ​​the AGV on the horizontal plane. By combining the AGV height data with the two-dimensional grid map, the three-dimensional spatial area covered by the AGV from the ground to the top of the vehicle during actual operation can be determined. This area is defined as the AGV's operating space.

[0027] Secondly, based on the obtained Automated Guided Vehicle (AGV) activity space, the target electromagnetic field spatiotemporal distribution model is extracted from the three-dimensional dynamic electromagnetic field spatiotemporal distribution model. Specifically, the three-dimensional dynamic electromagnetic field spatiotemporal distribution model describes the electromagnetic field distribution throughout the entire testing workshop from the ground to the ceiling. However, the operation and electromagnetic safety impact of the AGV are limited to the activity space occupied by its vehicle body. Therefore, it is necessary to spatially trim the established three-dimensional dynamic electromagnetic field spatiotemporal distribution model according to the obtained AGV activity space, retaining only the electromagnetic field distribution data within the AGV activity space, thereby obtaining a more concise target electromagnetic field spatiotemporal distribution model that focuses on the AGV's operating area.

[0028] Furthermore, the acquired spatiotemporal distribution model of the target electromagnetic field is mapped onto a two-dimensional grid map of the testing workshop to generate an electromagnetic situation heat map. The two-dimensional grid map of the testing workshop is a digital map in which the physical ground area of ​​the testing workshop is divided into uniform grids according to a set resolution, and each grid cell is assigned positional coordinate information.

[0029] Specifically, the target electromagnetic field spatiotemporal distribution model includes real-time field strength values ​​at various horizontal coordinate positions within the automated guided vehicle's (AGV) operational space. The real-time field strength value at each horizontal coordinate point in this model is assigned to a corresponding grid point on a two-dimensional grid map. Subsequently, using data visualization technology, different colors are assigned to each grid point based on its field strength value, for example, a gradient from blue representing low field strength to red representing high field strength. Finally, an electromagnetic situation heatmap is generated on the two-dimensional grid map, visually displaying the spatial distribution and dynamic changes of the electromagnetic field strength within the detection workshop, where each grid point corresponds to a quantified real-time field strength value.

[0030] This electromagnetic situational heatmap provides real-time electromagnetic safety constraint information accurate to each spatial grid for the path planning of automated guided vehicles, enabling path planning decisions to intuitively identify and avoid high field strength risk areas.

[0031] S30: Load the geometric state data, kinematic state data, and mission state data of the automated guided vehicle, and combine them with the electromagnetic situation heat map to construct a composite state space; Furthermore, the geometric, kinematic, and mission status data of the automated guided vehicle (AGV) are loaded. Specifically, during the AGV's transportation mission, its operating status is comprehensively affected by its own physical properties, dynamic motion characteristics, and the current mission objective. It is necessary to obtain information such as its current position, heading, speed, and the insulation level requirements of the transported equipment to comprehensively describe the AGV's condition at a given moment, thus providing a foundation for subsequent decision-making.

[0032] Furthermore, a composite state space is constructed by combining the aforementioned electromagnetic situation heatmap. This composite state space is a multi-dimensional information vector that deeply integrates the autonomous guided vehicle's own state information with external electromagnetic environment information. It is used to provide a comprehensive and integrated decision-making basis for the upper-level reinforcement learning policy network, enabling it to make path planning decisions that balance transportation efficiency and electromagnetic safety, while being aware of the vehicle's own condition and the risks of the surrounding electromagnetic environment.

[0033] Specifically, the geometric state data, kinematic state data, and mission state data of the automated guided vehicle are loaded, and combined with the electromagnetic situation heatmap, a composite state space is constructed, including: Load the current coordinates, heading angle, speed, target point coordinates, and static obstacle positions of the automated guided vehicle, and add them to the geometric state data; The linear velocity, angular velocity, and acceleration of the automated guided vehicle are loaded and added to the kinematic state data; Load the insulation level requirements of the current transportation equipment and add them to the task status data; The real-time field strength value of the current position of the automated guided vehicle, the field strength gradient on the predetermined trajectory in front of the automated guided vehicle, and the relative distance between the automated guided vehicle and the nearest high-voltage charged body are extracted from the electromagnetic situation heat map and added to the electromagnetic situation status data. The geometric state data, kinematic state data, task state data, and electromagnetic situational state data are combined to obtain the composite state space.

[0034] First, the current coordinates, heading angle, speed, target point coordinates, and static obstacle positions of the automated guided vehicle (AGV) are loaded and added to the geometric state data. Specifically, the AGV obtains its current coordinate position in a two-dimensional grid map through its onboard positioning system, acquires its current heading angle through its inertial measurement unit, and obtains its current speed through its motor encoder. Simultaneously, it reads the target point coordinates for this transport mission from the mission management system and loads the position information of static obstacles from a pre-stored environmental map. All of this data together constitutes the geometric state data describing the AGV's current spatial position, attitude, motion trend, and the relationship between the mission target and the static environment.

[0035] Secondly, the linear velocity, angular velocity, and acceleration of the automated guided vehicle (AGV) are added to the kinematic state data. Specifically, the AGV's motion control system reads the vehicle's linear velocity, angular velocity during turns, and current acceleration value in real time. This kinematic state data describes the AGV's motion capability and dynamic response characteristics in the next moment, providing a basis for the upper-level reinforcement learning policy network to plan actions that conform to the vehicle's physical motion laws.

[0036] Simultaneously, the insulation level requirements of the current transport equipment are loaded and added to the task status data. Specifically, the specific type of the transport equipment is parsed from the task instructions of the automated guided vehicle, and the corresponding insulation level requirements are queried. This insulation level requirement represents the upper limit of electromagnetic field strength that the equipment can safely withstand, and it is included as part of the task status data to ensure the safety of the transported goods during path planning.

[0037] In addition, the real-time field strength value of the current position of the automated guided vehicle, the field strength gradient on the predetermined trajectory in front of the automated guided vehicle, and the relative distance between the automated guided vehicle and the nearest high-voltage charged body are extracted from the electromagnetic situation heat map obtained above and added to the electromagnetic situation status data.

[0038] Specifically, firstly, based on the current coordinates of the automated guided vehicle (AGV), the real-time field strength value of the corresponding grid point is retrieved from the electromagnetic situation heatmap. Secondly, along the trajectory of the AGV's current course, the field strength values ​​of each grid point on the trajectory are extracted, and their rate of change is calculated as the field strength gradient to characterize the changing trend of the electromagnetic environment ahead. Finally, based on the distribution of high field strength regions in the electromagnetic situation heatmap, the distance between the AGV's current position and the region containing the nearest high-voltage charged body is calculated. These data collectively constitute electromagnetic situation status data describing the current electromagnetic environment risk of the AGV and its nearby changing trends.

[0039] Finally, the geometric state data, kinematic state data, task state data, and electromagnetic situational awareness data are merged to obtain a composite state space. Specifically, the data from these four aspects are vectorized and concatenated to form a high-dimensional state vector that comprehensively represents the Automated Guided Vehicle's current state, task constraints, and electromagnetic environment risks—the composite state space. This composite state space serves as the sole input to the subsequent upper-level reinforcement learning policy network for decision-making, ensuring that the decision-making process comprehensively considers all relevant factors.

[0040] S40: Based on the composite state space, a safety barrier reinforcement learning decision framework is constructed, which includes an upper-layer reinforcement learning policy network and a lower-layer safety barrier modifier. The upper-layer reinforcement learning policy network outputs a preliminary expected action, and the lower-layer safety barrier modifier performs real-time safety verification on the preliminary expected action based on an electromagnetic safety threshold. Furthermore, based on the obtained composite state space, a safety barrier reinforcement learning decision framework is constructed, comprising an upper-layer reinforcement learning policy network and a lower-layer safety barrier modifier. Specifically, the upper-layer reinforcement learning policy network is a neural network model built based on deep reinforcement learning algorithms. This network takes the composite state space as input and the initial desired action of the automated guided vehicle (AGV) as output, and is responsible for exploring and learning how to plan an efficient transportation path in a complex electromagnetic environment. The lower-layer safety barrier modifier is a real-time safety verification and correction module built based on a preset electromagnetic safety threshold. This module performs hard safety verification on the initial desired action of the upper layer and generates a safety correction action to ensure that the AGV does not violate the electromagnetic safety red line when the verification fails.

[0041] The safety barrier reinforcement learning decision framework it comprises is a hierarchical decision architecture that combines upper-level intelligent exploration and optimization with lower-level hard safety constraints. It is used to maximize the efficiency of transportation path planning while ensuring the absolute safety of automated guided vehicles.

[0042] Specifically, based on the aforementioned composite state space, a security barrier reinforcement learning decision framework is constructed, comprising an upper-layer reinforcement learning policy network and a lower-layer security barrier modifier, including: The upper-layer reinforcement learning policy network is constructed based on a deep reinforcement learning algorithm, wherein the input nodes of the upper-layer reinforcement learning policy network correspond to the composite state space, and the output nodes correspond to the initial expected action. The reward function of the upper-layer reinforcement learning policy network is set, wherein the reward function includes a first reward component representing the path length, a second reward component representing the transportation time, a third penalty component representing the autonomous guided vehicle entering a high field strength region, and a fourth penalty component representing the autonomous guided vehicle's sudden stop and turn. The safety barrier function of the underlying safety barrier modifier is set based on the electromagnetic safety threshold, wherein the safety barrier function includes a field strength safety threshold and a field strength change rate threshold; Connect the output node of the upper-layer reinforcement learning policy network to the input node of the lower-layer security barrier modifier to obtain the security barrier reinforcement learning decision framework.

[0043] First, a high-level reinforcement learning policy network is constructed based on a deep reinforcement learning algorithm. This high-level reinforcement learning policy network is constructed using a deep neural network structure, with the number of input layer nodes corresponding to the dimension of the composite state space, ensuring that all state information in the composite state space can be received by the network. The output layer nodes of the network correspond to various preliminary expected actions that the automated guided vehicle can perform, such as the expected changes in linear velocity and angular velocity. The function of this high-level reinforcement learning policy network is to calculate a preliminary expected action aimed at optimizing long-term cumulative rewards based on the input composite state through forward propagation.

[0044] Secondly, a reward function for the upper-layer reinforcement learning policy network is set. This reward function is a multi-objective weighted summation function used to guide the learning direction of the upper-layer reinforcement learning policy network. Specifically, the reward function includes four components: the first reward component is related to the path length traversed by the automated guided vehicle (AGV) from the starting point to the target point; the shorter the path, the higher the reward value. The second reward component is related to the time consumed in completing the transportation task; the shorter the time, the higher the reward value. The third penalty component is related to the AGV's behavior when entering high-field-strength areas. When the real-time field strength value at the current position of the AGV exceeds a preset safety threshold, a negative penalty value is applied to the reward function. This component is used to drive the upper-layer reinforcement learning policy network to actively avoid high-field-strength areas and ensure electromagnetic safety. The fourth penalty component is related to the AGV's driving stability. When the AGV's action output causes the rate of change of linear velocity or angular velocity to exceed a preset smoothing threshold, a negative penalty value is applied to the reward function. This component is used to suppress unstable driving behaviors such as sudden stops and sharp turns, ensuring the stability of the transportation equipment.

[0045] Furthermore, a safety barrier function for the underlying safety barrier corrector is set based on electromagnetic safety thresholds. The core of the underlying safety barrier corrector is a safety barrier function defined according to pre-set electromagnetic safety thresholds. Specifically, the safety barrier function includes two threshold parameters: the first is the field strength safety threshold, representing the upper limit of electromagnetic field strength that the automated guided vehicle (AGV) and its onboard equipment can safely withstand. This field strength safety threshold is set comprehensively based on electromagnetic field exposure safety standards, equipment insulation level requirements, and the electromagnetic immunity level of the AGV control system. The second is the field strength change rate threshold, representing the maximum rate of change of electromagnetic field strength that the AGV can accept per unit time. This field strength change rate threshold is set based on the response speed of the AGV control system, the lag time of the actuators, and the maximum adjustment capability of the motor driver, to avoid control interference caused by entering regions of drastic field strength changes.

[0046] The underlying security barrier modifier embeds a security barrier function built based on the aforementioned electromagnetic security threshold. This security barrier function serves as a hard security boundary constraint mechanism, used to verify the initial expected action output by the upper-layer reinforcement learning policy network in real time.

[0047] Furthermore, by connecting the output nodes of the upper-layer reinforcement learning policy network to the input nodes of the lower-layer safety barrier corrector, a safety barrier reinforcement learning decision framework is obtained. Specifically, the initial expected action output by the upper-layer reinforcement learning policy network is used as input and fed into the lower-layer safety barrier corrector. The lower-layer safety barrier corrector performs real-time safety verification on the initial expected action based on its safety barrier function: if the predicted field strength value or field strength change rate after executing the initial expected action does not exceed the corresponding threshold, the safety barrier function determines that the action is safe and allows the action to be output; if the predicted value exceeds any threshold, the safety barrier function determines that the action is illegal and automatically calculates a minimum intervention correction action that can forcibly pull the automated guided vehicle back to a safe state to replace the original initial expected action. This forms a hierarchical decision framework in which the upper-layer network pursues optimal path efficiency, and the lower-layer barrier ensures absolute operational safety.

[0048] S50: When the underlying security barrier modifier determines that the preliminary expected action passes the security check, it outputs the preliminary expected action as the actual control command. Specifically, when the bottom-level safety barrier modifier determines that the preliminary expected action has passed the safety check, it outputs the preliminary expected action as the actual control command. The purpose of this step is to ensure that the preliminary expected action output by the upper-level reinforcement learning policy network after exploration and optimization, and verified as safe by the bottom-level safety barrier modifier, can be executed smoothly. This ensures the optimal path planning efficiency pursued by the upper-level policy network while guaranteeing the absolute safety of the automated guided vehicle operation.

[0049] Specifically, this step directly translates the initial expected actions that have passed the hard safety check into actual control commands, driving the automated guided vehicle to travel along the predetermined planned path, thereby achieving a balance between safety and efficiency.

[0050] Specifically, when the underlying security barrier modifier determines that the preliminary expected action passes the security check, it outputs the preliminary expected action as an actual control command, including: The initial desired action is loaded from the underlying security barrier modifier; Based on the initial expected action, predict the first position of the automated guided vehicle at the next moment; Query the first predicted field strength value corresponding to the first position from the electromagnetic situation heat map; When the first predicted field strength value is less than the field strength safety threshold and the field strength change rate is less than the field strength change rate threshold, the preliminary expected action is determined to pass the safety check. The initial desired action is output as the actual control command.

[0051] First, the initial expected action is loaded from the bottom-level safety barrier modifier. The bottom-level safety barrier modifier receives and temporarily stores the initial expected action output by the upper-level reinforcement learning policy network. This initial expected action includes information such as the expected changes in linear velocity and angular velocity.

[0052] Secondly, based on the initial expected action, the first position of the automated guided vehicle (AGV) at the next moment is predicted. Specifically, by combining the current geometric and kinematic state data of the AGV, such as its coordinates, heading angle, and velocity, the control command represented by the initial expected action is applied to the vehicle's kinematic model. Through dead reckoning or dynamic simulation calculations, the spatial position that the AGV will reach at the next moment after executing the initial expected action is predicted, and this position is defined as the first position.

[0053] Then, the first predicted field strength value corresponding to the first location is retrieved from the electromagnetic situation heatmap. Specifically, based on the predicted planar coordinates of the first location, a location query is performed in the real-time updated electromagnetic situation heatmap, and the real-time field strength value corresponding to that coordinate grid point is extracted. This value is defined as the first predicted field strength value. This first predicted field strength value is used to characterize the electromagnetic field strength that the automated guided vehicle will face in the next moment if the initial expected action is performed.

[0054] Specifically, when the first predicted field strength value is less than the field strength safety threshold, and the field strength change rate is less than the field strength change rate threshold, the preliminary expected action is deemed to have passed the safety check. The first predicted field strength value obtained from the query is compared with the preset field strength safety threshold, and the field strength change rate from the current position to the first field strength position is calculated and compared with the preset field strength change rate threshold. Only when the first predicted field strength value is strictly less than the field strength safety threshold, and the field strength change rate is also strictly less than the field strength change rate threshold, does the underlying safety barrier corrector determine that the preliminary expected action complies with the electromagnetic safety hard constraints and passes the check.

[0055] After the safety verification is passed, the underlying safety barrier corrector forwards the initial expected action directly to the underlying actuator of the automated guided vehicle as an actual control command, driving the vehicle to move according to the action command, thereby achieving optimal path planning efficiency under the premise of safety.

[0056] S60: When the underlying security barrier modifier determines that the initial expected action has failed the security check, it generates a minimum intervention security correction action, uses the security correction action as the actual control instruction, and feeds back the event of the security check failure as a penalty to the upper-layer reinforcement learning policy network for learning. Furthermore, when the underlying safety barrier corrector determines that the initial expected action has failed the safety check, it means that if the initial expected action output by the upper-layer reinforcement learning policy network is executed, it will cause the automated guided vehicle to enter a dangerous area where the field strength value exceeds the safety threshold or the field strength change rate is too large, which may lead to equipment failure or safety accidents.

[0057] At this point, a safety correction action with minimal intervention needs to be generated. This action serves as the actual control command to ensure that the automated guided vehicle (AGV) can be forcibly pulled back to a safe trajectory, avoiding contact with the electromagnetic safety red line. Simultaneously, events that fail this safety check are fed back as penalties to the upper-level reinforcement learning policy network for learning. This encourages the network to proactively avoid states and actions that could trigger safety barriers in subsequent decisions, thereby continuously optimizing its strategy and gradually achieving the goal of improving path planning efficiency without sacrificing safety.

[0058] Specifically, when the underlying security barrier modifier determines that the initial expected action has failed the security check, it generates a minimum intervention security correction action, uses this security correction action as the actual control command, and simultaneously feeds the event of this security check failure as a penalty term back to the upper-layer reinforcement learning policy network for learning, including: The initial desired action is loaded from the underlying security barrier modifier; Based on the initial expected action, predict the second position of the automated guided vehicle at the next moment; Query the second predicted field strength value corresponding to the second position from the electromagnetic situation heat map; When the second predicted field strength value is greater than or equal to the field strength safety threshold, or the field strength change rate is greater than or equal to the field strength change rate threshold, it is determined that the preliminary expected action has failed the safety check. The underlying safety barrier corrector calculates a minimum intervention correction action to pull the automated guided vehicle back to a safe trajectory, sets it as the safety correction action, and outputs the actual control command. The event that fails the security check is used as a penalty and fed back to the upper-layer reinforcement learning policy network to update the network parameters of the upper-layer reinforcement learning policy network.

[0059] First, the initial expected action is loaded from the bottom-level safety barrier modifier. The bottom-level safety barrier modifier receives and temporarily stores the initial expected action output by the upper-level reinforcement learning policy network. This initial expected action includes information such as the expected changes in linear velocity and angular velocity.

[0060] Secondly, based on the initial expected action, the second position of the automated guided vehicle (AGV) at the next moment is predicted. Specifically, by combining the current geometric and kinematic state data of the AGV, such as its coordinates, heading angle, and velocity, the control command represented by the initial expected action is applied to the vehicle's kinematic model. Through dead reckoning or dynamic simulation calculations, the spatial position that the AGV will reach at the next moment after executing the initial expected action is predicted, and this position is defined as the second position.

[0061] Then, the second predicted field strength value corresponding to the second position is retrieved from the electromagnetic situation heat map. Based on the predicted planar coordinates of the second position, a location query is performed in the real-time updated electromagnetic situation heat map to extract the real-time field strength value corresponding to the grid point of that coordinate. This value is defined as the second predicted field strength value, which is used to characterize the electromagnetic field strength that the automated guided vehicle will face in the next moment if the initial expected action is performed.

[0062] Specifically, when the second predicted field strength value is greater than or equal to the field strength safety threshold, or the field strength change rate is greater than or equal to the field strength change rate threshold, the preliminary expected action is determined to have failed the safety check. The obtained second predicted field strength value is compared with the preset field strength safety threshold, and the field strength change rate from the current position to the second field strength position is calculated and compared with the preset field strength change rate threshold. As long as either the second predicted field strength value is greater than or equal to the field strength safety threshold, or the field strength change rate is greater than or equal to the field strength change rate threshold, the underlying safety barrier corrector determines that the preliminary expected action violates the electromagnetic safety hard constraints and fails the check.

[0063] Furthermore, the underlying safety barrier corrector calculates a minimum intervention correction action to pull the automated guided vehicle (AGV) back to a safe trajectory, designated as the safety correction action, and outputs the actual control command. Specifically, after a verification failure, the underlying safety barrier corrector activates a safety correction mechanism. This mechanism solves an optimization problem based on the safety barrier function and the current state. The goal of this optimization problem is to find, among all actions that would enable the AGV to meet the electromagnetic safety threshold in the next time step, the action with the smallest deviation from the initial expected action, which is designated as the minimum intervention correction action. This minimum intervention correction action is the action with the smallest deviation from the initial expected action output by the upper-layer reinforcement learning policy network, obtained by solving the optimization problem among all candidate actions that would enable the AGV to meet the electromagnetic safety threshold in the next time step.

[0064] The minimal intervention correction action is designated as the safety correction action and directly output to the underlying actuator of the automated guided vehicle (AGV) as an actual control command, ensuring that the vehicle is forcibly guided to a safe trajectory and avoids entering dangerous electromagnetic areas. Simultaneously, the event of failing this safety check is treated as a penalty and fed back to the upper-level reinforcement learning policy network to update its network parameters. The underlying safety barrier corrector records the event in which the safety barrier is triggered and transmits this event as a clear negative feedback signal, i.e., a penalty, to the upper-level reinforcement learning policy network. Upon receiving this penalty, the upper-level reinforcement learning policy network treats it as an unfavorable decision-making experience for subsequent network parameter updates and learning optimization. This allows it to tend to output the initially desired action that will not trigger the safety barrier in similar complex state spaces, thus gradually learning to autonomously avoid high-risk actions over a long learning process.

[0065] S70: Execute the actual control command to complete the path planning of the automated guided vehicle.

[0066] Finally, the actual control commands are executed to complete the path planning for the automated guided vehicle (AGV). This step involves sending the actual control commands, which are ultimately generated in the safety barrier reinforcement learning decision framework, including the preliminary expected actions or safety correction actions that have undergone safety verification, to the underlying actuators of the AGV. These are then translated into physical movements such as wheel rotation and speed adjustment, thereby driving the AGV to travel along the planned path.

[0067] The purpose of this step is to translate the decision-making results into actual vehicle control behavior. By continuously executing each control command, the automated guided vehicle can travel safely and efficiently from the starting point to the target point, ultimately completing the path planning and execution closed loop of the entire transportation task.

[0068] In addition, the method also includes: A simulation environment synchronized with the physical testing workshop is constructed in the digital twin system, wherein the simulation environment includes the three-dimensional dynamic electromagnetic field spatiotemporal distribution model; The security barrier reinforcement learning decision framework is pre-trained offline in the simulation environment to obtain pre-trained model parameters. The pre-trained model parameters are deployed to the controller of the physical automated guided vehicle; During the actual operation of the physically automated guided vehicle, the parameters of the pre-trained model are fine-tuned online based on the real-time collected electromagnetic situational awareness data.

[0069] Furthermore, as a preferred embodiment of the present invention, the method further includes the following steps: A simulation environment synchronized with the physical testing workshop is constructed within the digital twin system. This simulation environment includes a three-dimensional dynamic spatiotemporal distribution model of electromagnetic fields. The digital twin system is an integrated simulation platform encompassing multiple physics fields, scales, and probabilities. Utilizing the physical model of the testing workshop, sensor data, and historical operational data, this platform can create a digital mirror image in virtual space that perfectly matches the geometry, equipment layout, and operational status of the real physical testing workshop.

[0070] By importing the three-dimensional dynamic electromagnetic field spatiotemporal distribution model constructed in the aforementioned steps into the digital twin system, the simulation environment can realistically reproduce the dynamic changes of the electromagnetic field in the physical workshop, providing a highly realistic and risk-free virtual test field for subsequent model training.

[0071] Secondly, the safety barrier reinforcement learning decision framework is pre-trained offline in a simulation environment to obtain pre-trained model parameters. Specifically, the constructed safety barrier reinforcement learning decision framework is deployed in a digital twin simulation environment. This environment allows for the rapid generation of large amounts of training data and enables the policy network to undergo extensive trial-and-error learning, thus providing sufficient offline training for the upper-layer reinforcement learning policy network. During the training process, the safety barrier modifier also plays a role in the simulation environment, ensuring that the exploration process always complies with electromagnetic safety constraints. Through extensive iterative learning, the policy network can converge to a relatively optimal level, and the network parameters at this point are the pre-trained model parameters.

[0072] Specifically, this offline pre-training step can give the model a good initial performance before applying it to a real vehicle, reducing the risks and debugging time in the online learning phase on the real vehicle.

[0073] Then, the pre-trained model parameters are deployed to the controller of the physical automated guided vehicle (AGV). Once the offline pre-training achieves the desired results, the trained pre-trained model parameters are exported from the digital twin system and loaded into the controller of the physical AGV. This controller, as the computational core of the AGV, uses the loaded pre-trained model parameters as the initial weights for its upper-layer reinforcement learning policy network, enabling the physical vehicle to possess a basic and reasonable path planning capability from the start of actual operation.

[0074] Meanwhile, during the actual operation of the physically-guided transport vehicle (AGV), the parameters of the pre-trained model are fine-tuned online based on real-time collected electromagnetic situational awareness data. Specifically, after the AGV is put into actual operation, its onboard mobile field strength monitoring terminal and the fixed sensor array in the workshop continuously collect real-time electromagnetic situational awareness data for further online fine-tuning of the deployed pre-trained model parameters. Since there may be subtle differences between the simulation environment and the physical world, online fine-tuning allows the model parameters to adaptively adjust to better match the electromagnetic dynamic characteristics of real-world working conditions and the motion response characteristics of the physical vehicle, thereby achieving continuous optimization and adaptive improvement of model performance.

[0075] In summary, this preferred embodiment, by combining digital twin pre-training with online fine-tuning, can solve the adaptability problem of reinforcement learning from simulation to reality.

[0076] In addition, the method also includes: Obtain the historical disaster recovery records of the testing workshop, wherein the historical disaster recovery records include records of automated guided vehicles losing control due to electromagnetic interference and records of events where the safe distance between the automated guided vehicles and high-voltage live conductors is insufficient; Based on the historical disaster recovery records, the number of times that automated guided vehicles (AGV) lose control events and insufficient safety distance events occurred in several two-dimensional grid areas were statistically analyzed. Two-dimensional grid regions with a frequency greater than or equal to a preset threshold are extracted from the aforementioned historical occurrences and configured as high-risk warning regions; When the planned route of the automated guided vehicle passes through the high-risk warning area, route replanning is performed.

[0077] Meanwhile, as another preferred embodiment of the present invention, the method further includes the following steps: First, the historical disaster recovery records of the testing workshop are obtained. These records include those documenting Automated Guided Vehicle (AGV) loss of control events caused by electromagnetic interference and those documenting insufficient safe distances between AGVs and high-voltage live conductors. The historical disaster recovery records are a collection of data extracted and compiled from the testing workshop's operation logs, accident reports, and maintenance records. These records detail specific information about AGV control system malfunctions, communication interruptions, or navigation deviations caused by strong electromagnetic interference during a historical period, as well as information on AGVs violating safe distance regulations and approaching high-voltage live conductors excessively. The historical disaster recovery records typically include key data such as the time and location of the event, the environmental parameters at the time, and the consequences of the event.

[0078] Secondly, based on historical disaster recovery records, the number of times each automated guided vehicle (AGV) out-of-control event and insufficient safety distance event occurred in several two-dimensional grid areas was statistically analyzed. Specifically, the two-dimensional grid map of the testing workshop was used as the statistical benchmark. For each event in the historical disaster recovery records, it was categorized into the corresponding grid area on the map based on its location coordinates at the time of occurrence. After all historical events were categorized, the number of AGV out-of-control events and insufficient safety distance events occurring in each grid area were counted, forming a statistical value of the historical occurrence count for each grid area.

[0079] Then, extract a number of two-dimensional grid areas from the historical occurrences that are greater than or equal to a preset threshold number, and configure them as high-risk warning areas.

[0080] First, a preset frequency threshold is set to define whether a grid area belongs to a high-risk zone with frequent accidents. This preset frequency threshold is comprehensively set based on the historical operating cycle length of the testing workshop, the overall accident incidence rate, and safety management level requirements; for example, it can be set to 5 times. The historical occurrence statistics of all grid areas are compared one by one with this preset frequency threshold to extract grid areas where the number of out-of-control events or insufficient safety distance events reaches or exceeds the preset frequency threshold. All extracted grid areas are specially marked in the system and configured as high-risk warning areas, indicating that the area has experienced multiple safety accidents in the past and has a high potential risk.

[0081] Specifically, when the planned route of the automated guided vehicle (AGV) passes through a high-risk warning area, route replanning is performed. During the AGV's route planning process, it checks whether the currently planned route intersects with any of the configured high-risk warning areas. Once it is detected that the planned route will cross or enter any high-risk warning area, the route replanning mechanism will be triggered immediately.

[0082] Specifically, this route replanning is a dynamic adjustment process designed to ensure that automated guided vehicles (AGVs) safely avoid high-risk warning areas while minimizing the impact on their original transport missions. Optionally, this can be achieved through the following methods: First, local route adjustment, where the main body of the original global route remains unchanged, but only the local road segments overlapping with the high-risk warning area are replanned to generate a detour route to avoid the area; second, speed planning adjustment, where when a complete detour is not possible, the AGV slows down before entering the high-risk warning area, passing through smoothly at a lower speed, while simultaneously strengthening operational status monitoring to address potential risks; third, global route replanning, where when the high-risk warning area covers a large area and cannot be effectively avoided through local adjustments, a globally optimal route that completely avoids all high-risk warning areas is searched based on the current starting point, target point, and an updated risk map.

[0083] The aforementioned path replanning mechanism can reduce the operational risks of automated guided vehicles in areas with a history of high accident rates, thereby improving the overall safety and reliability of transportation operations.

[0084] In summary, this preferred embodiment utilizes historical accident data to identify inherent high-risk areas and proactively avoids them during the path planning stage, thereby reducing the likelihood of similar accidents recurring from the source and improving the disaster recovery capability and long-term operational safety of the automated guided vehicle in complex electromagnetic environments.

[0085] In summary, the embodiments of this application have at least the following technical effects: First, this invention utilizes a fixed sensor array and an automated guided vehicle (AGV) to collaboratively collect electromagnetic situational data, constructing a three-dimensional dynamic spatiotemporal distribution model of the electromagnetic field. This model is then overlaid on a two-dimensional grid map to form an electromagnetic situational heatmap, achieving high-precision real-time perception of the electromagnetic environment in the inspection workshop. Second, this invention constructs a composite state space integrating vehicle geometry, kinematics, task status, and real-time electromagnetic situational awareness, providing a comprehensive input of multi-dimensional safety factors for reinforcement learning decision-making. Simultaneously, this invention proposes a safety barrier reinforcement learning decision-making framework comprising an upper-layer reinforcement learning policy network and a lower-layer safety barrier modifier. The upper-layer policy network explores the optimal path to improve transportation efficiency, while the lower-layer safety barrier modifier performs hard safety checks on preliminary desired actions based on electromagnetic safety thresholds, ensuring that the AGV always operates within the electromagnetic safety boundary. Furthermore, this invention feeds back events that fail safety checks as penalties to the upper-layer reinforcement learning policy network for learning, prompting the policy network to proactively avoid dangerous actions.

[0086] Ultimately, this invention solves the problem of balancing transportation efficiency and electromagnetic safety in the path planning of automated guided vehicles under high-voltage electrical testing environments, and improves the safety and intelligence level of autonomous transportation operations in complex environments.

[0087] Example 2, as Figure 2 As shown, based on the same inventive concept as the reinforcement learning-based AGV transportation path planning method provided in Embodiment 1, this embodiment of the invention also provides a reinforcement learning-based AGV transportation path planning system, including: The data loading module 11 is used to load electromagnetic situational awareness data from the testing workshop, wherein the electromagnetic situational awareness data includes first field strength data collected by a fixed electromagnetic field sensor array and second field strength data collected by an automated guided vehicle. The electromagnetic situation modeling module 12 is used to construct a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the testing workshop based on the electromagnetic situation perception data, and obtain an electromagnetic situation heat map superimposed on a two-dimensional grid map, wherein each grid point in the electromagnetic situation heat map corresponds to a real-time field strength value. The composite state space construction module 13 is used to load the geometric state data, kinematic state data and mission state data of the automated guided vehicle, and combine them with the electromagnetic situation heat map to construct a composite state space; The decision framework construction module 14 is used to construct a safety barrier reinforcement learning decision framework based on the composite state space, which includes an upper-layer reinforcement learning policy network and a lower-layer safety barrier modifier. The upper-layer reinforcement learning policy network outputs a preliminary expected action, and the lower-layer safety barrier modifier performs real-time safety verification on the preliminary expected action based on an electromagnetic safety threshold. The first instruction output module 15 is used to output the preliminary expected action as an actual control instruction when the underlying security barrier modifier determines that the preliminary expected action has passed the security check. The second instruction output and feedback module 16 is used to generate a minimum intervention safety correction action when the underlying security barrier corrector determines that the initial expected action has failed the security check, and to use the safety correction action as the actual control instruction. At the same time, the event of failing the security check is fed back to the upper-layer reinforcement learning policy network as a penalty item for learning. The path planning execution module 17 is used to execute the actual control instructions and complete the path planning of the automated guided vehicle.

[0088] Specifically, the data loading module 11 is used for: Specifically, the electromagnetic situational awareness data from the testing workshop is loaded, including: The fixed electromagnetic field sensor array is loaded with the first field strength data collected at a preset sampling frequency; The second field strength data collected during the operation of the automated guided vehicle is loaded by the mobile field strength monitoring terminal mounted on the automated guided vehicle. The first field strength data and the second field strength data are aggregated to the central processor via a wireless communication network and added to the electromagnetic situational awareness data.

[0089] The electromagnetic situation modeling module 12 is specifically used for: Specifically, based on the electromagnetic situational awareness data, a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the detection workshop is constructed to obtain an electromagnetic situational heat map superimposed on a two-dimensional grid map, including: Based on the principle of finite element analysis, a basic model of the static electromagnetic field distribution of the high-voltage electrical testing equipment in the testing workshop is established. Based on the electromagnetic situational awareness data, the static electromagnetic field distribution model is dynamically corrected by a data fusion algorithm to obtain the three-dimensional dynamic electromagnetic field spatiotemporal distribution model. Based on the height data of the automated guided vehicle and the two-dimensional grid map, the activity space of the automated guided vehicle is obtained; Based on the activity space of the automated guided vehicle, the spatiotemporal distribution model of the target electromagnetic field is extracted from the three-dimensional dynamic electromagnetic field spatiotemporal distribution model; The spatiotemporal distribution model of the target electromagnetic field is mapped onto a two-dimensional grid map of the testing workshop to generate the electromagnetic situation heat map, wherein each grid point in the electromagnetic situation heat map corresponds to a real-time field strength value.

[0090] The composite state space construction module 13 is specifically used for: Specifically, the geometric state data, kinematic state data, and mission state data of the automated guided vehicle are loaded, and combined with the electromagnetic situation heatmap, a composite state space is constructed, including: Load the current coordinates, heading angle, speed, target point coordinates, and static obstacle positions of the automated guided vehicle, and add them to the geometric state data; The linear velocity, angular velocity, and acceleration of the automated guided vehicle are loaded and added to the kinematic state data; Load the insulation level requirements of the current transportation equipment and add them to the task status data; The real-time field strength value of the current position of the automated guided vehicle, the field strength gradient on the predetermined trajectory in front of the automated guided vehicle, and the relative distance between the automated guided vehicle and the nearest high-voltage charged body are extracted from the electromagnetic situation heat map and added to the electromagnetic situation status data. The geometric state data, kinematic state data, task state data, and electromagnetic situational state data are combined to obtain the composite state space.

[0091] The decision framework construction module 14 is specifically used for: Specifically, based on the aforementioned composite state space, a security barrier reinforcement learning decision framework is constructed, comprising an upper-layer reinforcement learning policy network and a lower-layer security barrier modifier, including: The upper-layer reinforcement learning policy network is constructed based on a deep reinforcement learning algorithm, wherein the input nodes of the upper-layer reinforcement learning policy network correspond to the composite state space, and the output nodes correspond to the initial expected action. The reward function of the upper-layer reinforcement learning policy network is set, wherein the reward function includes a first reward component representing the path length, a second reward component representing the transportation time, a third penalty component representing the autonomous guided vehicle entering a high field strength region, and a fourth penalty component representing the autonomous guided vehicle's sudden stop and turn. The safety barrier function of the underlying safety barrier modifier is set based on the electromagnetic safety threshold, wherein the safety barrier function includes a field strength safety threshold and a field strength change rate threshold; Connect the output node of the upper-layer reinforcement learning policy network to the input node of the lower-layer security barrier modifier to obtain the security barrier reinforcement learning decision framework.

[0092] The first instruction output module 15 is specifically used for: Specifically, when the underlying security barrier modifier determines that the preliminary expected action passes the security check, it outputs the preliminary expected action as an actual control command, including: The initial desired action is loaded from the underlying security barrier modifier; Based on the initial expected action, predict the first position of the automated guided vehicle at the next moment; Query the first predicted field strength value corresponding to the first position from the electromagnetic situation heat map; When the first predicted field strength value is less than the field strength safety threshold and the field strength change rate is less than the field strength change rate threshold, the preliminary expected action is determined to pass the safety check. The initial desired action is output as the actual control command.

[0093] The second instruction output and feedback module 16 is specifically used for: Specifically, when the underlying security barrier modifier determines that the initial expected action has failed the security check, it generates a minimum intervention security correction action, uses this security correction action as the actual control command, and simultaneously feeds the event of this security check failure as a penalty term back to the upper-layer reinforcement learning policy network for learning, including: The initial desired action is loaded from the underlying security barrier modifier; Based on the initial expected action, predict the second position of the automated guided vehicle at the next moment; Query the second predicted field strength value corresponding to the second position from the electromagnetic situation heat map; When the second predicted field strength value is greater than or equal to the field strength safety threshold, or the field strength change rate is greater than or equal to the field strength change rate threshold, it is determined that the preliminary expected action has failed the safety check. The underlying safety barrier corrector calculates a minimum intervention correction action to pull the automated guided vehicle back to a safe trajectory, sets it as the safety correction action, and outputs the actual control command. The event that fails the security check is used as a penalty and fed back to the upper-layer reinforcement learning policy network to update the network parameters of the upper-layer reinforcement learning policy network.

[0094] Specifically, the path planning execution module 17 is used for: The actual control commands are executed to complete the path planning for the automated guided vehicle.

[0095] In addition, the method also includes: A simulation environment synchronized with the physical testing workshop is constructed in the digital twin system, wherein the simulation environment includes the three-dimensional dynamic electromagnetic field spatiotemporal distribution model; The security barrier reinforcement learning decision framework is pre-trained offline in the simulation environment to obtain pre-trained model parameters. The pre-trained model parameters are deployed to the controller of the physical automated guided vehicle; During the actual operation of the physically automated guided vehicle, the parameters of the pre-trained model are fine-tuned online based on the real-time collected electromagnetic situational awareness data.

[0096] It also includes: Obtain the historical disaster recovery records of the testing workshop, wherein the historical disaster recovery records include records of automated guided vehicles losing control due to electromagnetic interference and records of events where the safe distance between the automated guided vehicles and high-voltage live conductors is insufficient; Based on the historical disaster recovery records, the number of times that automated guided vehicles (AGV) lose control events and insufficient safety distance events occurred in several two-dimensional grid areas were statistically analyzed. Two-dimensional grid regions with a frequency greater than or equal to a preset threshold are extracted from the aforementioned historical occurrences and configured as high-risk warning regions; When the planned route of the automated guided vehicle passes through the high-risk warning area, route replanning is performed.

Claims

1. An AGV transportation path planning method based on reinforcement learning, characterized in that, include: The electromagnetic situational awareness data of the testing workshop is loaded, wherein the electromagnetic situational awareness data includes first field strength data collected by a fixed electromagnetic field sensor array and second field strength data collected by an automated guided vehicle. Based on the electromagnetic situational awareness data, a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the testing workshop is constructed to obtain an electromagnetic situational heat map superimposed on a two-dimensional grid map, wherein each grid point in the electromagnetic situational heat map corresponds to a real-time field strength value. Load the geometric state data, kinematic state data, and mission state data of the automated guided vehicle, and combine them with the electromagnetic situation heat map to construct a composite state space; Based on the composite state space, a safety barrier reinforcement learning decision framework is constructed, which includes an upper-layer reinforcement learning policy network and a lower-layer safety barrier modifier. The upper-layer reinforcement learning policy network outputs a preliminary expected action, and the lower-layer safety barrier modifier performs real-time safety verification on the preliminary expected action based on an electromagnetic safety threshold. When the underlying security barrier modifier determines that the preliminary expected action passes the security check, it outputs the preliminary expected action as the actual control command. When the underlying security barrier modifier determines that the initial expected action has failed the security check, it generates a minimum intervention security correction action, uses the security correction action as the actual control instruction, and feeds back the event of this security check failure as a penalty to the upper-layer reinforcement learning policy network for learning. The actual control commands are executed to complete the path planning for the automated guided vehicle.

2. The AGV transportation path planning method based on reinforcement learning according to claim 1, characterized in that, Load electromagnetic situational awareness data from the testing workshop, including: The fixed electromagnetic field sensor array is loaded with the first field strength data collected at a preset sampling frequency; The second field strength data collected during the operation of the automated guided vehicle is loaded by the mobile field strength monitoring terminal mounted on the automated guided vehicle. The first field strength data and the second field strength data are aggregated to the central processor via a wireless communication network and added to the electromagnetic situational awareness data.

3. The AGV transportation path planning method based on reinforcement learning according to claim 1, characterized in that, Based on the electromagnetic situational awareness data, a three-dimensional dynamic spatiotemporal distribution model of the electromagnetic field in the testing workshop is constructed to obtain an electromagnetic situational heat map superimposed on a two-dimensional grid map, including: Based on the principle of finite element analysis, a basic model of the static electromagnetic field distribution of the high-voltage electrical testing equipment in the testing workshop is established. Based on the electromagnetic situational awareness data, the static electromagnetic field distribution model is dynamically corrected by a data fusion algorithm to obtain the three-dimensional dynamic electromagnetic field spatiotemporal distribution model. Based on the height data of the automated guided vehicle and the two-dimensional grid map, the activity space of the automated guided vehicle is obtained; Based on the activity space of the automated guided vehicle, the spatiotemporal distribution model of the target electromagnetic field is extracted from the three-dimensional dynamic electromagnetic field spatiotemporal distribution model; The spatiotemporal distribution model of the target electromagnetic field is mapped onto a two-dimensional grid map of the testing workshop to generate the electromagnetic situation heat map, wherein each grid point in the electromagnetic situation heat map corresponds to a real-time field strength value.

4. The AGV transportation path planning method based on reinforcement learning according to claim 1, characterized in that, Load the geometric state data, kinematic state data, and mission state data of the automated guided vehicle, and combine them with the electromagnetic situation heatmap to construct a composite state space, including: Load the current coordinates, heading angle, speed, target point coordinates, and static obstacle positions of the automated guided vehicle, and add them to the geometric state data; The linear velocity, angular velocity, and acceleration of the automated guided vehicle are loaded and added to the kinematic state data; Load the insulation level requirements of the current transportation equipment and add them to the task status data; The real-time field strength value of the current position of the automated guided vehicle, the field strength gradient on the predetermined trajectory in front of the automated guided vehicle, and the relative distance between the automated guided vehicle and the nearest high-voltage charged body are extracted from the electromagnetic situation heat map and added to the electromagnetic situation status data. The geometric state data, kinematic state data, task state data, and electromagnetic situational state data are combined to obtain the composite state space.

5. The AGV transportation path planning method based on reinforcement learning according to claim 1, characterized in that, Based on the aforementioned composite state space, a security barrier reinforcement learning decision framework is constructed, comprising an upper-layer reinforcement learning policy network and a lower-layer security barrier modifier, including: The upper-layer reinforcement learning policy network is constructed based on a deep reinforcement learning algorithm, wherein the input nodes of the upper-layer reinforcement learning policy network correspond to the composite state space, and the output nodes correspond to the initial expected action. The reward function of the upper-layer reinforcement learning policy network is set, wherein the reward function includes a first reward component representing the path length, a second reward component representing the transportation time, a third penalty component representing the autonomous guided vehicle entering a high field strength region, and a fourth penalty component representing the autonomous guided vehicle's sudden stop and turn. The safety barrier function of the underlying safety barrier modifier is set based on the electromagnetic safety threshold, wherein the safety barrier function includes a field strength safety threshold and a field strength change rate threshold; Connect the output node of the upper-layer reinforcement learning policy network to the input node of the lower-layer security barrier modifier to obtain the security barrier reinforcement learning decision framework.

6. The AGV transportation path planning method based on reinforcement learning according to claim 5, characterized in that, When the underlying security barrier modifier determines that the preliminary expected action passes the security check, it outputs the preliminary expected action as the actual control command, including: The initial desired action is loaded from the underlying security barrier modifier; Based on the initial expected action, predict the first position of the automated guided vehicle at the next moment; Query the first predicted field strength value corresponding to the first position from the electromagnetic situation heat map; When the first predicted field strength value is less than the field strength safety threshold and the field strength change rate is less than the field strength change rate threshold, the preliminary expected action is determined to pass the safety check. The initial desired action is output as the actual control command.

7. The AGV transportation path planning method based on reinforcement learning according to claim 5, characterized in that, When the underlying security barrier modifier determines that the initial expected action fails the security check, it generates a minimum intervention security correction action, uses this security correction action as the actual control command, and simultaneously feeds the event of failing the security check as a penalty term back to the upper-layer reinforcement learning policy network for learning, including: The initial desired action is loaded from the underlying security barrier modifier; Based on the initial expected action, predict the second position of the automated guided vehicle at the next moment; Query the second predicted field strength value corresponding to the second position from the electromagnetic situation heat map; When the second predicted field strength value is greater than or equal to the field strength safety threshold, or the field strength change rate is greater than or equal to the field strength change rate threshold, it is determined that the preliminary expected action has failed the safety check. The underlying safety barrier corrector calculates a minimum intervention correction action to pull the automated guided vehicle back to a safe trajectory, sets it as the safety correction action, and outputs the actual control command. The event that fails the security check is used as a penalty and fed back to the upper-layer reinforcement learning policy network to update the network parameters of the upper-layer reinforcement learning policy network.

8. The AGV transportation path planning method based on reinforcement learning according to claim 1, characterized in that, Also includes: A simulation environment synchronized with the physical testing workshop is constructed in the digital twin system, wherein the simulation environment includes the three-dimensional dynamic electromagnetic field spatiotemporal distribution model; The security barrier reinforcement learning decision framework is pre-trained offline in the simulation environment to obtain pre-trained model parameters. The pre-trained model parameters are deployed to the controller of the physical automated guided vehicle; During the actual operation of the physically automated guided vehicle, the parameters of the pre-trained model are fine-tuned online based on the real-time collected electromagnetic situational awareness data.

9. The AGV transportation path planning method based on reinforcement learning according to claim 1, characterized in that, Also includes: Obtain the historical disaster recovery records of the testing workshop, wherein the historical disaster recovery records include records of automated guided vehicles losing control due to electromagnetic interference and records of events where the safe distance between the automated guided vehicles and high-voltage live conductors is insufficient; Based on the historical disaster recovery records, the number of times that automated guided vehicles (AGV) lose control events and insufficient safety distance events occurred in several two-dimensional grid areas were statistically analyzed. Two-dimensional grid regions with a frequency greater than or equal to a preset threshold are extracted from the aforementioned historical occurrences and configured as high-risk warning regions; When the planned route of the automated guided vehicle passes through the high-risk warning area, route replanning is performed.

10. An AGV transportation path planning system based on reinforcement learning, characterized in that, The method for executing the reinforcement learning-based AGV transportation path planning method according to any one of claims 1-9 includes: The data loading module is used to load electromagnetic situational awareness data from the testing workshop. The electromagnetic situational awareness data includes first field strength data collected by a fixed electromagnetic field sensor array and second field strength data collected by an automated guided vehicle. The electromagnetic situation modeling module is used to construct a three-dimensional dynamic electromagnetic field spatiotemporal distribution model of the testing workshop based on the electromagnetic situation perception data, and obtain an electromagnetic situation heat map superimposed on a two-dimensional grid map, wherein each grid point in the electromagnetic situation heat map corresponds to a real-time field strength value. The composite state space construction module is used to load the geometric state data, kinematic state data and mission state data of the automated guided vehicle, and combine them with the electromagnetic situation heat map to construct a composite state space; The decision framework construction module is used to construct a safety barrier reinforcement learning decision framework based on the composite state space, which includes an upper-layer reinforcement learning policy network and a lower-layer safety barrier modifier. The upper-layer reinforcement learning policy network outputs a preliminary expected action, and the lower-layer safety barrier modifier performs real-time safety verification on the preliminary expected action based on an electromagnetic safety threshold. The first instruction output module is used to output the preliminary expected action as an actual control instruction when the underlying security barrier modifier determines that the preliminary expected action has passed the security check. The second instruction output and feedback module is used to generate a minimum intervention safety correction action when the underlying security barrier corrector determines that the initial expected action has failed the security check. The safety correction action is used as the actual control instruction, and the event of failing the security check is fed back to the upper-layer reinforcement learning policy network as a penalty for learning. The path planning execution module is used to execute the actual control commands and complete the path planning of the automated guided vehicle.