Communication perception control integrated networked formation control method and device

By constructing a dynamic obstacle avoidance risk map and optimizing the communication cycle, combined with ISCC global information, a unified decision-making framework for multi-agent formation control was realized. This solved the problems of poor communication and control performance, collision risk, and limited scale in multi-agent networked control, and achieved efficient and safe formation control.

CN121541691APending Publication Date: 2026-02-17BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511660809.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing multi-agent networked control schemes suffer from poor communication and control performance under heavy communication in large-scale intelligent machines, extreme collision risks and unstable formations in dynamic environments, and are limited in scale and have excessive communication overhead.

Method used

A dynamic obstacle avoidance risk map based on the conditional risk value model is constructed, the upper bound of the communication cycle is determined, and the global information of the integrated communication, perception, and control (ISCC) is combined with the information. The communication cycle allocation is optimized through a multi-agent formation control algorithm to realize a unified decision-making framework for perception and control.

Benefits of technology

It solves the problem of poor performance caused by busy communication in large-scale intelligent machines, the risk of extreme collisions and unstable formations in dynamic environments, reduces communication overhead, and achieves safe, efficient and scalable formation control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121541691A_ABST
    Figure CN121541691A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a communication perception control integrated networked formation control method and device. The method comprises the following steps: constructing a dynamic obstacle avoidance risk map based on a conditional value-at-risk model; determining an upper bound of a communication period according to a stability condition of a trajectory optimization equation of the AGV formation, and determining a communication period distribution mechanism of the AGV formation according to the upper bound of the communication period and the dynamic obstacle avoidance risk map; and based on the dynamic obstacle avoidance risk map and the communication period distribution mechanism, in combination with ISCC global information, determining a multi-agent formation control algorithm based on the ISCC, and according to the multi-agent formation control algorithm of the ISCC, performing networked formation control on the AGV formation. According to the scheme of the application, the problems of poor communication and control performance, extreme collision risk, unstable formation, limited multi-agent scale and overlarge communication overhead in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of integrated communication, sensing, and control technology, specifically to a networked formation control method and apparatus that integrates communication, sensing, and control. Background Technology

[0002] With the rapid development of artificial intelligence and communication technologies, multi-agent systems, represented by multi-vehicle cooperative control, have been widely applied in fields such as industrial automation and smart logistics. However, behind their large-scale successful commercial applications, multi-agent systems rely on pre-set information such as indoor simultaneous localization and mapping (SLAM) maps and radio frequency identification (RFID) tags to assist local perception for positioning and navigation. They often move along predetermined static trajectories and lack the ability to accurately perceive and adapt to unknown dynamic environments. Especially in the highly dynamic and complex scenarios of flexible automotive assembly workshops, existing control systems with low communication and weak geometric constraints can only support simple low-speed formation control tasks, making it difficult to meet the complex task requirements of ultra-high precision formation synchronization and smooth and safe transportation for multiple agents in the workshop. Therefore, there is an urgent need to build an efficient, scalable, and smooth multi-agent networked control formation system, including two aspects: perception-based formation control and communication-based formation control methods. The aim is to design accurate perception and positioning methods and low-latency, scalable communication network topologies to improve the obstacle avoidance capability and cooperative control performance of the control formation system.

[0003] Existing perception-based formation control schemes often employ isolated perception processes and lack efficient communication sharing strategies. This results in limited perception information from a single agent being unable to be fully utilized and shared by the multi-agent control system, thus restricting control performance. This also means that the aforementioned low-speed, simple formation control strategies are unsuitable for highly dynamic environments, especially in narrow corridors or densely packed material areas—complex environments where dynamic and static obstacles change in real time. Due to limited perception range and accuracy, delayed obstacle avoidance braking exacerbates the risk of congestion and collisions.

[0004] Existing communication-based formation control schemes suffer from several drawbacks in practical applications. Distributed formation control struggles to acquire global environmental information, exhibits poor adaptability to dynamic and complex environments, suffers from significant synchronization errors, and is unsuitable for maintaining strict geometric formations. Centralized formation control, limited by base station communication resources and computing power, struggles to support large-scale agent-based formation collaboration. As node size increases and task complexity rises, frequent information exchange and calibration significantly increase communication overhead, leading to substantial increases in closed-loop latency and synchronization errors, making it difficult to meet the demands of ultra-low latency and highly efficient collaborative formation tasks.

[0005] The two types of formation control methods described above, which use dynamic models, rely on pre-set rules for control decisions. This means that formation control cannot guarantee the effectiveness of decisions in complex and unknown environments. Furthermore, current research still suffers from the limitations of independent redundancy designs for perception, communication, and control. Specifically, firstly, the system calculates the optimal control strategy based solely on local perception information such as relative position and velocity. This often fails to achieve efficient collaborative control due to limitations in communication resources and underutilization of perception information. Particularly under conditions of limited communication resources, this independent redundancy design hinders the allocation of multi-dimensional resources across perception, communication, and control. It struggles to balance the conflicting relationship between the communication overhead of global perception information and the accumulation of perception errors, often sacrificing potential trajectory smoothness to improve agent maneuverability. This leads to a series of disjointed emergency decelerations and accelerations during obstacle avoidance and navigation, increasing synchronization errors and potentially causing material crushing, deformation, or even drop damage. Furthermore, due to insufficient systematic research on the coupling relationship between communication, perception, and control modules in high-precision cooperative control of multiple agents, the increasing user-defined optimization objectives and the number of controllable agents make it difficult for the system to effectively coordinate multi-module interactions and dynamically adapt to complex task requirements. This leads to a surge in the complexity of the formation control problem, making finding an effective solution extremely challenging. Simply meeting basic requirements, such as ensuring safe obstacle avoidance and formation in static obstacle environments while minimizing travel time, is already a challenge. Achieving high maneuverability and strict geometric formation synchronization in dynamic and complex environments is even more difficult. This is why traditional formation control research has struggled to move from structured, known static obstacle environments to unpredictable, complex, dynamic, and flexible manufacturing workshop environments.

[0006] In summary, existing technical solutions have the following problems: poor communication and control performance due to busy communication of large-scale intelligent machines in traditional multi-agent networked control schemes; extreme collision risks and formation instability in dynamic environments; and limited multi-agent scale and excessive communication overhead. Summary of the Invention

[0007] At least one embodiment of this application provides a networked formation control method and apparatus that integrates communication, sensing, and control, to solve the problems of poor communication and control performance caused by busy communication of large-scale intelligent machines in traditional multi-agent networked control schemes in the prior art, the extreme collision risk and formation instability in dynamic environments, and the problems of limited multi-agent scale and excessive communication overhead.

[0008] To solve the above-mentioned technical problems, this application is implemented as follows:

[0009] In a first aspect, embodiments of this application provide a networked formation control method integrating communication, sensing, and control, including:

[0010] Construct a dynamic obstacle avoidance risk map based on the conditional value-at-risk model;

[0011] Based on the stability condition of the trajectory optimization equation of the AGV formation, the upper bound of the communication cycle is determined, and based on the upper bound of the communication cycle and the dynamic obstacle avoidance risk map, the communication cycle allocation mechanism of the AGV formation is determined.

[0012] Based on the dynamic obstacle avoidance risk map and the communication cycle allocation mechanism, and combined with the ISCC global information of the integrated communication, perception and control system, a multi-agent formation control algorithm based on ISCC is determined.

[0013] The AGV formation is networked and controlled according to the ISCC multi-agent formation control algorithm.

[0014] Secondly, embodiments of this application provide a networked formation control device integrating communication, sensing, and control, comprising:

[0015] The first building module is used to construct a dynamic obstacle avoidance risk map based on the conditional value at risk model;

[0016] The first determining module is used to determine the upper bound of the communication cycle based on the stability condition of the trajectory optimization equation of the automated guided vehicle (AGV) formation, and to determine the communication cycle allocation mechanism of the AGV formation based on the upper bound of the communication cycle and the dynamic obstacle avoidance risk map.

[0017] The second determining module is used to determine the ISCC-based multi-agent formation control algorithm based on the dynamic obstacle avoidance risk map and the communication cycle allocation mechanism, combined with the ISCC global information of the integrated communication perception and control.

[0018] The first processing module is used to perform networked formation control on the AGV formation according to the ISCC multi-agent formation control algorithm.

[0019] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in any of the first aspects.

[0020] Fourthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the method described in any of the first aspects.

[0021] Compared with existing technologies, the networked formation control method and apparatus integrating communication, perception, and control provided in this application determines the upper bound of the communication cycle based on the stability of the AGV formation trajectory optimization equation, and then allocates the cycle on demand using a dynamic obstacle avoidance risk map to balance communication load and control performance, thus solving the problem of poor communication and control performance caused by busy communication in large-scale intelligent machines. The dynamic obstacle avoidance risk map perceives collision risks in real time, and the stability conditions of trajectory optimization ensure the formation basis. After determination, the ISCC multi-agent formation control algorithm is used to integrate risk and communication decisions. In case of extreme risks, high-frequency communication and emergency trajectory adjustment are triggered to solve the problems of extreme collision risks and formation instability in dynamic environments. Based on the ISCC multi-agent formation control algorithm, the AGV formation is networked and controlled. The communication cycle is allocated on demand to reduce the total overhead, and a unified decision framework is provided to adapt to large-scale formations, breaking through the scale limitation and solving the problems of limited multi-agent scale and excessive communication overhead. Attached Figure Description

[0022] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0023] Figure 1 A flowchart of the networked formation control method integrating communication, sensing, and control provided in the embodiments of this application;

[0024] Figure 2 A schematic diagram of the scene structure of the dynamic obstacle avoidance risk map provided in the embodiments of this application;

[0025] Figure 3 This is a schematic diagram illustrating the implementation process of the networked formation control method integrating communication, sensing, and control as described in this application.

[0026] Figure 4 This is a schematic diagram illustrating the variation of formation synchronization accuracy of different algorithms in a dynamic and complex environment, as shown in this application.

[0027] Figure 5 A schematic diagram illustrating the trajectory change pattern of a multi-agent formation in a complex environment with alternating dynamic and static obstacles, as shown in this application;

[0028] Figure 6 This is a structural diagram of a networked formation control device integrating communication, sensing, and control, provided in an embodiment of this application. Detailed Implementation

[0029] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, without limiting the number of objects; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0030] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc.; an indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.

[0031] To enable those skilled in the art to better understand the embodiments of this application, the following description is provided first:

[0032] This application can be applied to the field of Integrated Sensing Communication and Control (ISCC), which refers to the unified design of communication, sensing and control functions through the joint design of linear / nonlinear controllers, industrial wireless network protocols, networked collaborative sensing and other means. This enables the wireless network to achieve high-precision and refined sensing functions, as well as real-time and efficient control performance while conducting high-quality communication interaction, thereby improving the overall performance and business capabilities of the manufacturing workshop control system.

[0033] This application relates to the field of multi-agent formation control, which refers to the need for multiple agents (such as automated guided vehicles and robots) to form a stable formation through communication, perception, and control coordination in complex dynamic environments such as flexible manufacturing. This formation aims to maintain formation and avoid collisions while collaboratively completing production tasks such as material handling and assembly. Unlike single-agent independent operation, multi-agent formation control emphasizes real-time interaction and task division among the group. It must satisfy individual kinematic and dynamic constraints while addressing challenges such as network communication latency, information asymmetry, and environmental uncertainty.

[0034] This application can be applied to the field of closed-loop control, which refers to the highly spatially and temporally differentiated closed-loop control business in intelligent manufacturing workshops, such as intelligent machine control closed loop and production line full life cycle management closed loop. This process involves a large number of communication, sensing and control processes between sensing devices, edge servers and control devices, requiring flexible and efficient closed-loop control.

[0035] This application can be applied to the field of wireless network control, which refers to combining wireless communication and automatic control technologies to use wireless networks to achieve real-time monitoring and control of production lines, robots and automated guided vehicles, thereby improving the production efficiency of intelligent manufacturing.

[0036] This application provides a networked formation control method and apparatus that integrates communication, sensing, and control. The method and apparatus are based on the same concept, and since they solve problems based on similar principles, their implementations can be mutually referenced; repeated details will not be elaborated further.

[0037] This application proposes a networked formation control method and apparatus integrating communication, sensing, and control. To accurately characterize the correlation between communication, sensing, and control in multi-agent networked control processes, this scheme establishes a novel multi-agent formation control system model based on the integration of communication, sensing, and control. Based on this, a dynamic obstacle avoidance risk map based on the Conditional Value at Risk (CVaR) model is constructed to quantify the potential congestion and collision risks of dynamic obstacles sensed by the formation under the influence of network latency. This allows sufficient time for smooth speed changes and formation synchronization, without sacrificing movement speed to reduce the risk of material compression and deformation during emergency braking and startup. Subsequently, this scheme models the formation control problem as a partially observable Markov decision process and proposes a Multi-Agent Proximal Policy Optimization (MAPPO) strategy based on global ISCC states. From the perspective of communication and perception-assisted enhanced control, to reduce the additional overhead of ISCC state interaction, this strategy designs a dynamic communication cycle allocation mechanism based on global state. The AGV dynamically adjusts its interaction cycle with the base station based on ISCC state observations, balancing synchronization accuracy and communication overhead. Furthermore, this strategy employs a heterogeneous actor-critic framework to alleviate gradient conflicts and improve the control efficiency of heterogeneous agents, thereby improving the number of controllable AGVs and travel time. This provides a safe, efficient, and scalable formation control paradigm for industrial collaborative handling scenarios.

[0038] For further details, please refer to Figure 1 This application provides a networked formation control method integrating communication, sensing, and control, comprising:

[0039] Step 11: Construct a dynamic obstacle avoidance risk map based on the conditional value at risk model.

[0040] This application constructs a dynamic obstacle avoidance risk map based on the Conditional Value at Risk (CVaR) model. CVaR is a statistical model that quantifies extreme risks. Compared to traditional risk assessments (such as variance), it focuses more on low-probability, high-loss collision scenarios, such as sudden AGV deflections and dynamic obstacle intrusions. Step 11 collects real-time data on the AGV formation's position, speed, and environmental obstacle dynamics. Combined with the CVaR model, it calculates the collision risk probability and loss severity in different areas and at different times, ultimately generating a dynamically updated risk map. Areas with higher risk values ​​represent greater extreme collision risks. This step 11 overcomes the limitations of traditional static risk assessment, enabling risk to be perceived, quantified, and dynamically tracked, providing a risk benchmark for subsequent communication resource allocation and control decisions.

[0041] Step 12: Determine the upper bound of the communication cycle based on the stability condition of the trajectory optimization equation of the AGV formation, and determine the communication cycle allocation mechanism of the AGV formation based on the upper bound of the communication cycle and the dynamic obstacle avoidance risk map.

[0042] Here, step 12 determines the upper bound of the communication cycle and constructs a communication cycle allocation mechanism. Determining the upper bound of the communication cycle in the trajectory optimization equation of the AGV formation is the core of ensuring formation stability. Its stability condition has strict constraints on the communication cycle. If the communication cycle is too long, the delay in information exchange between AGVs will lead to lag in control commands, thereby disrupting formation synchronization. Step 12 determines the upper bound of the communication cycle by solving the stability condition of the trajectory optimization equation, thus avoiding formation instability caused by communication delays at the root. The communication cycle allocation mechanism, combined with the dynamic obstacle avoidance risk map generated in step 11, allocates the communication cycle as needed. For example, in high-risk areas, short communication cycles are allocated, allowing AGVs to exchange position and speed information frequently, ensuring real-time response to control commands and avoiding collisions; in low-risk areas, long communication cycles are allocated to reduce redundant information transmission and lower communication load.

[0043] Step 13: Based on the dynamic obstacle avoidance risk map and the communication cycle allocation mechanism, and combined with the ISCC global information of the integrated communication, perception and control system, determine the ISCC-based multi-agent formation control algorithm.

[0044] This application integrates the dynamic obstacle avoidance risk information from step 11 with the communication cycle allocation rules from step 12 to construct a multi-agent formation control algorithm based on ISCC. The functions of this ISCC-based multi-agent formation control algorithm are: to quantify the correlation between risk, communication, and control; and to provide a unified decision-making framework for AGV formation, enabling the model to output the optimal control strategy regardless of the number of AGVs or changes in environmental risks.

[0045] Step 14: Perform networked formation control on the AGV formation according to the ISCC multi-agent formation control algorithm.

[0046] In this embodiment, the ISCC-based multi-agent formation control algorithm in step 13 is used as a control command generator to output specific control schemes in real time: If the model identifies a sudden increase in the CVaR value of a certain area (extreme collision risk): a short communication cycle is triggered, instructing the AGVs in that area to decelerate and adjust their course, while simultaneously notifying surrounding AGVs to maintain a safe distance; if the model identifies an expansion of the formation size but low overall risk: resources are allocated according to a long communication cycle, instructing newly added AGVs to connect to the existing communication network to maintain formation synchronization without additional communication overhead. Ultimately, through a closed loop of perception-decision-control, the dynamic and stable operation of the AGV formation is achieved.

[0047] In this embodiment, step 12 allocates communication cycles on demand, enabling high-frequency communication only in high-risk areas and reducing communication frequency in low-risk areas, thus reducing redundant communication from the source. Step 13 uses an ISCC-based multi-agent formation control algorithm to ensure that communication resources are prioritized for AGVs with high control requirements, avoiding waste of channel resources. Ultimately, this achieves a match between communication load and control needs, preventing congestion while ensuring real-time control commands. Step 11 uses a CVaR model specifically to capture low-probability, high-loss collision scenarios, and a dynamic obstacle avoidance risk map updates risk locations in real time, allowing the system to detect threats in advance. Step 14, based on model commands, triggers short communication cycles and emergency obstacle avoidance control in extreme risk situations to avoid collisions. For unstable formations, step 12 uses the stability condition of the trajectory optimization equation to determine the upper bound of the communication cycle, ensuring that communication delays do not disrupt formation synchronization. Step 14 generates control commands based on the ISCC multi-agent formation control algorithm, balancing obstacle avoidance and formation maintenance to prevent formation collapse due to obstacle avoidance. To address the issue of excessive communication overhead, step 12's on-demand allocation of communication cycles ensures that communication overhead does not increase linearly with the number of AGVs, but rather dynamically adjusts with the number of risk areas, maintaining overall control even in large-scale scenarios. For situations where the scale of multiple agents is limited, step 13's ISCC-based multi-agent formation control algorithm provides a general decision-making framework. Regardless of whether the number of AGVs is 10 or 100, the model can adapt to risks and communication resources without requiring separate control logic design for different scales. Simultaneously, the communication cycle allocation mechanism in step 12 supports seamless integration of new AGVs (allocating cycles based on risk, without needing to reconstruct the communication network), breaking through the scale limitations of traditional solutions.

[0048] Reference Figure 2 As shown, this application takes collaborative material handling in a flexible automotive assembly workshop as a typical scenario and establishes a novel integrated multi-agent formation control system model that combines sensing and control. This model is helpful for studying the formation control process of multiple AGVs in industrial environments. Based on this, a dynamic obstacle avoidance risk map based on CVaR is constructed to quantify the potential collision risk of dynamic obstacles, enhance the potential sensing range and accuracy, ensure smooth speed changes, and reduce the risk of material compression deformation without sacrificing speed. Secondly, formation control is modeled as a partially observable Markov decision process, and a MAPPO formation control strategy based on ISCC global information is proposed. A dynamic communication cycle allocation mechanism is introduced to balance the communication overhead of global sensing information and the accumulation of sensing errors, thereby reducing communication costs, improving control efficiency, and increasing the number of controllable AGVs. Furthermore, by decoupling the actor-critic framework into different agent networks, this application improves the training efficiency of heterogeneous agents and reduces gradient conflicts, thereby improving the number of controllable AGVs and travel time, providing a safe, efficient, and scalable formation control paradigm for industrial collaborative material handling scenarios. Further details can be found by referring to... Figure 3 As shown, this application presents the MAPPO formation control strategy flow based on ISCC global information.

[0049] To assess the impact of dynamic environmental uncertainties on formation control, this application constructs a dynamic obstacle avoidance risk map based on CVaR from the perspectives of communication and perception-assisted enhanced control. By correlating multi-source perception data with parameters such as pose control information and communication latency, and considering stricter tail expectations to constrain extreme risks such as congestion collisions, the severity of potential congestion collisions between AGVs and dynamic obstacles during closed-loop formation control is quantified. Sufficient reaction time to potential environmental conditions is reserved to smoothly adjust speed, reducing the risk of material compression, friction, deformation, or even drop damage caused by emergency acceleration or braking, thus ensuring smooth and accurate formation synchronization.

[0050] To achieve efficient and safe movement of multi-agent formations in complex dynamic environments, this application transforms the formation problem into a partially observable Markov decision process model and proposes a multi-agent formation control algorithm based on ISCC (ISCC-MAPPO). In ISCC-MAPPO, leaders and followers, as heterogeneous agents, can be trained to simultaneously handle optimal control and communication decisions in continuous and discrete action spaces. During training, to reduce the additional communication overhead caused by ISCC global information, this application proposes a dynamic communication cycle allocation mechanism based on network control state, ensuring high-precision formation synchronization while reducing communication overhead and latency. Furthermore, to promote efficient cooperation among heterogeneous agents, a heterogeneous actor-critic framework is proposed for both leader and follower agents, thereby improving the performance and applicability of the ISCC-MAPPO algorithm and increasing the scale of controllable agents.

[0051] Furthermore, step 11 above includes:

[0052] Collect interaction parameters between the current AGV formation and dynamic obstacles; the dynamic obstacles are used to represent AGVs in neighboring formations of the current AGV formation.

[0053] Based on the interaction parameters, under the boundary condition that the AGV in the current AGV formation and the dynamic obstacle simultaneously decelerate at maximum speed and collide just when they come to a stop, the relative speed and safe braking distance between the AGV and the dynamic obstacle are calculated.

[0054] Based on the relative speed and the safe braking distance, determine the congestion collision probability expression;

[0055] Based on the congestion collision probability expression, determine the objective function of the dynamic obstacle avoidance risk map;

[0056] A dynamic obstacle avoidance risk map is constructed based on N sets of collision probability samples and the objective function.

[0057] In this embodiment, considering that when AGVs navigate through complex and rapidly changing dense obstacles on production lines, the high closed-loop latency and limited environmental perception range of the AGVs themselves make them prone to low-probability but high-loss congestion collisions with dynamic and static obstacles, causing material to fall and be damaged. Although traditional methods sacrifice agent motion performance, using low-speed travel to compensate for a longer reaction distance to avoid collisions, they still rely on a safe distance from obstacles within a local perception range to make obstacle avoidance decisions. This inevitably leads to emergency braking decisions, and drastic speed fluctuations still increase synchronization errors, exacerbating the shift in the center of gravity of materials being transported in the formation, resulting in squeezing, friction, deformation, and even falling damage. Therefore, in addition to considering a more comprehensive perception range for early warning and collision avoidance, it is also necessary to consider the smoothness of speed changes to ensure strict formation synchronization and material safety. To more accurately assess the impact of dynamic environmental uncertainties on multi-agent obstacle avoidance and formation synchronization, this application designs a method to quantify the potential congestion and collision risks of perceived AGV formations and dynamic obstacles under the influence of network latency. A dynamic obstacle avoidance risk map based on CVaR is established, allowing sufficient time for smooth speed changes and formation synchronization, without sacrificing movement speed to reduce the risk of material compression and deformation during emergency braking and startup.

[0058] Specifically, by assessing the potential collision loss intensity of obstacle interactions in the worst-case scenario rather than simply the probability, CVaR predicts and quantifies the degree of motion conflict loss of AGV formations in the high-confidence interval of the risk map, upgrading traditional obstacle avoidance based on relative distance to risk value management under the collision probability distribution. This means that the dynamic obstacle avoidance risk map based on CVaR not only inherits the ability of probabilistic risk assessment to characterize environmental uncertainties, but also strengthens local path planning for low-probability, high-risk events such as material compression deformation and untimely emergency braking through the α-quantile tail risk focusing mechanism. This helps the motion space output smooth speed change strategies within the safety threshold, ensuring formation synchronization accuracy while maintaining stable driving.

[0059] In real-world industrial environments, multiple AGVs travel sequentially in formation along a circular production line. AGVs in neighboring formations can be considered dynamic obstacles with uncertainties. Since AGVs and other dynamic obstacles (such as AGVs in other formations) only begin braking and deceleration after receiving braking commands or global status information, it is necessary to calculate the distance traveled during the communication delay. Furthermore, considering the relative motion between the dynamic obstacles and the current AGV, the distance change should be based on their relative velocities. Assuming the worst-case scenario, the dynamic obstacles and... (This refers to any AGV in the AGV platoon, where n is the AGV number) They simultaneously decelerate at maximum speed and still collide upon coming to a complete stop. The relative speed at this point is... The expression for the first and safe braking distance The second expression is shown below:

[0060] Expression 1;

[0061] Expression 2;

[0062] in, The average network closed-loop delay of agents within the formation, and They are respectively and Current driving speed, and They are respectively and Current direction of travel and They are respectively and The maximum acceleration. Since static obstacles are usually in a fixed position and do not move, the AGV has already avoided static obstacles when planning its travel path, and can achieve obstacle avoidance by relying on its own sensors. In addition, vehicles often travel slowly in narrow corridors or areas with dense materials, and dynamic obstacles are relatively dense, indirectly calculating that the risk value of such areas is higher than that of open areas with sparse obstacles. If the current relative distance If the congestion occurs, a collision will happen. The expression for the probability of a congestion collision is: Based on the stochastic modeling of relative velocity, random variables are defined. Then the congestion collision probability is equivalent to expression three:

[0063] Expression three;

[0064] Then, in order to ensure that the tail risk of the collision probability does not exceed the safety threshold. Based on the expression for congestion collision probability, the objective function of the dynamic obstacle avoidance risk map is determined, i.e., the objective function CVaR based on the value at risk of congestion collision probability is designed. ,in The confidence level is given. Under the influence of network latency, the perceived risk of potential congestion collisions between AGV formations and dynamic obstacles can be quantified as a risk value. This can be represented by expression four.

[0065] Expression four;

[0066] in, It is the probability density function of the collision probability. This represents the maximum probability of an AGV colliding with an obstacle at a certain confidence level. Furthermore, N sets of samples are generated using Monte Carlo simulation to estimate the distribution of collision probabilities, which are then sorted in ascending order, and the first set is extracted. Collision probability data as The estimated value.

[0067] Finally, a two-dimensional map of the environment is constructed using real-time sensor and communication network feature information. Each grid area on the map is associated with a risk value, forming a dynamic obstacle avoidance risk map. The grid size is autonomously adjusted according to actual response time and speed smoothness requirements. By correlating perceived data with parameters such as dynamic pose information and network latency, and providing real-time global dynamic map feedback, the AGV can quickly respond to environmental changes. For example, it can slow down in advance to pass through high-risk areas, reserving time for potential collisions and reactions, or it can move at high speed through low-risk areas, significantly reducing speed fluctuations during obstacle avoidance and synchronization, ensuring smooth and accurate formation synchronization and material safety. At the same time, fully perceiving the risk value information of low-risk areas improves the average travel speed and reduces the overall travel time. In addition, using the dynamic obstacle avoidance risk map as input helps improve the accuracy of the reward mechanism in the deep learning network, improves the applicability of the learning algorithm, and meets the needs of extremely high-safety control scenarios, such as formation control, by more strictly constraining extreme risks through tail expectation.

[0068] Optionally, step 12 above includes:

[0069] Based on the motion process of AGV formations moving sequentially along a circular production line and a pre-set dynamic model, dynamic trajectory optimization equations are constructed.

[0070] Based on the stability conditions of the trajectory optimization equation, the AGV pose information, and the dynamic obstacle avoidance risk map, the safe distance and optimal speed parameters under the business cycle are calculated, and the upper bound of the communication cycle for maintaining the stability of the formation control system is determined. The stability conditions are used to determine that the following AGV can track the trajectory of the leading AGV, so that the position error and speed error within the formation converge to zero or are limited to a safe threshold.

[0071] In this application, when studying the sequential formation of multiple AGVs along a circular production line, this motion process can also be approximated as a Leader-Follower convoy dynamic model. Based on this, and using the dynamic trajectory equations of multiple AGVs proposed in previous work based on the influence of network closed-loop delay and safety distance, the smoothly moving AGVs can be guided by receiving state observation data. By analyzing information such as the pose of the intelligent agent and the dynamic obstacle avoidance risk map, parameters such as safe distance and optimal speed under typical business cycles are calculated. This leads to the derivation of the upper bound of the communication cycle for closed-loop interaction that maintains the stability of the control system. This can be represented by expression five:

[0072] Expression five;

[0073] ;

[0074] in, This represents the reference end-to-end latency in the 3GPP standard when an AGV performs industrial tasks with different cycle times. for Safe distance. For the optimal velocity function, . It is a cofactor, used to calculate and The remainder between these parameters ensures the communication cycle of the actual AGV. It is an integer multiple of the typical business cycle to maintain communication clock synchronization. This indicates the current communication cycle. This helps AGVs dynamically adjust their communication cycles, reducing the number of interactions between AGVs operating smoothly in open areas and base stations, improving network resource utilization, and reducing average network closed-loop latency.

[0075] When the closed-loop control delay is much smaller than the communication cycle, the agent's control commands have not yet been updated, and its motion state continues to execute the control decisions of the previous cycle within the current cycle. The agent's state update expression at this time is:

[0076] .

[0077] in Represents time t pose information, For rotation matrix, Represents time t Control commands.

[0078] Optionally, based on the upper bound of the communication cycle and the dynamic obstacle avoidance risk map, a communication cycle allocation mechanism for the AGV formation is determined, including:

[0079] According to the dynamic obstacle avoidance risk map, if the current area is a non-assembly area, it is determined to be a low-risk area; if the current area is an assembly area, it is determined to be a high-risk area.

[0080] In the low-risk area, and within the upper limit of the communication cycle, the communication cycle is adaptively increased, and the communication cycle is determined to be an integer multiple of the AGV control business cycle by using the cofunction; when the closed-loop control delay is less than the communication cycle, the AGV continues to execute the control decision of the previous cycle according to the updated motion state.

[0081] In the case of the high-risk area, the communication cycle is adjusted to a preset communication cycle, and a global risk map and high-precision pose information are obtained through high-frequency base station interaction.

[0082] In the embodiments of this application, during the determination of the dynamic communication cycle allocation mechanism based on networked control state, it is considered that the perception and synchronization errors in the pose information of the agent, which relies on local perception, will continue to accumulate over time. (where n represents any AGV in the AGV formation) requires a fixed interaction period to exchange high-precision multi-sensor data within the workshop with the base station to calibrate information and eliminate its own sensor errors. However, as the period of AGV's interaction with the base station on global state information decreases, while reducing perception and synchronization errors, it also introduces additional communication overhead. To balance the contradictory relationship between the global information communication interaction period and the accumulation of synchronization errors, this application proposes a dynamic communication period allocation mechanism based on networked control state. In this mechanism, the AGV dynamically outputs an appropriate communication interaction period within its action space based on ISCC state observations, such as motion state and congestion collision risk values ​​of the driving area. This ensures high-precision synchronization of formation control while significantly reducing communication overhead and closed-loop latency, and increasing the number of controllable AGVs per unit area.

[0083] Among them, the communication interaction cycle change process of multi-agent formation can be analyzed in two main ways.

[0084] Scenario 1: When multi-agent formations operate in areas with low CvaR risk values, such as non-assembly areas or areas with low AGV density, the control system tends to remain stable and is less susceptible to interference. Under traditional fixed interaction cycle mechanisms, although the base station receives sensing data from these stable AGVs multiple times, the calculated control commands are similar. However, the base station still ignores the impact of the control system's stable state on the communication network, continuing to transmit similar redundant control commands to the AGVs at a high frequency with a fixed 4ms interaction cycle. This results in a significant waste of network resources, as the communication network cannot handle the communication overhead of frequent interactions between large-scale intelligent machines. Therefore, in this scenario, this application uses the upper bound of the communication cycle as state space data to train the AGV control strategy and communication decisions in the Critic network. Specifically, this application can learn to calculate the optimal communication cycle for interacting control commands while ensuring stable formation control efficiency. This means that, under the premise of control system stability, this application can adaptively increase the communication cycle, reduce the number of times each agent accesses the base station and the communication overhead, thereby expanding the number of AGVs that the base station can control and increasing production efficiency.

[0085] Scenario 2: When multi-agent formations travel in areas with high CvaR risk values, such as assembly areas or high AGV density areas, the control system trajectory is easily affected by environmental changes. Since AGVs still rapidly accumulate pose and synchronization errors based on their own observations of local physical motion information, in areas with high CvaR risk values, AGVs rely more heavily on the dynamic obstacle avoidance risk map transmitted from the base station, high-precision pose information, and low-latency control decisions to avoid potential congestion and collision risks, calibrate their trajectories and pose information, and synchronize their formation. Therefore, in this scenario, the closed-loop control cycle in the AGV training results approximates the cycle time of the control operation, i.e. Strict formation and safe formation control should be given priority.

[0086] ;

[0087] Furthermore, because AGVs cannot receive global information from the base station during communication cycle intervals, such as... Because of this information, the interaction period reward function in the AGV's local Actor-Critic network remains unchanged, and thus the interaction period in the continuous action space calculated locally within the period interval is approximately a constant value.

[0088] In summary, by utilizing ISCC global state information, the proposed communication mechanism can dynamically adjust the AGV's communication cycle as needed, filtering out repetitive or low-value control commands, reducing the frequency of AGV access to the base station, and avoiding network congestion and high closed-loop latency caused by a large number of accesses. Furthermore, stable low-latency communication effectively reduces suboptimal decisions caused by lagging global state perception data, supports more accurate global state estimation, thereby enhancing the accuracy of the Critic network's value function evaluation and optimizing the policy's convergence performance. Subsequent simulations also demonstrate that the mechanism possesses stronger generalization capabilities; the dynamic communication cycle transmission mechanism supports training and execution in scenarios with a larger number of intelligent machines, making the policy more applicable in complex and changing environments.

[0089] Optionally, based on the dynamic obstacle avoidance risk map and the communication cycle allocation mechanism, and combined with the integrated ISCC global information, a multi-agent formation control algorithm based on ISCC is determined, including:

[0090] Determine the AGV formation control scenario and the type of intelligent agent; the type of intelligent agent includes lead AGV and follower AGV;

[0091] Determine the global state space and action space; the global state space is the integrated ISCC global information of communication, perception and control determined based on AGV pose, driving speed, formation parameters, the upper bound of the communication cycle, and the dynamic obstacle avoidance risk map;

[0092] Based on the global state space and the action space, a state transition function is constructed; the state transition function takes the current state and the current action as input and the probability distribution of the next state as output;

[0093] The aforementioned agent type is constructed, and a differentiated reward function is constructed. The differentiated reward function includes a leader AGV reward function and a follower AGV reward function. The leader AGV reward function includes obstacle avoidance, navigation, and corresponding interaction rewards for the leader AGV. The follower AGV reward function includes obstacle avoidance, formation, and corresponding interaction rewards for the follower AGV.

[0094] The execution-reward Actor-Critic network corresponding to the multi-agent reinforcement learning algorithm is split into a leading AGV Actor network, a following AGV Actor network, and a shared Critic network to construct a centralized training and decentralized execution CTDE framework.

[0095] Based on the differentiated reward function, the communication cycle allocation mechanism, and the CTDE framework, a multi-agent formation control algorithm based on ISCC is determined.

[0096] In this embodiment, the multi-agent formation control problem is formulated as a partially observable Markov Decision Process (POMDP) ​​under communication, perception, and control constraints. This step involves clarifying the scenario, defining core elements, designing differentiated mechanisms and frameworks, and constructing a POMDP decision model adapted to complex dynamic environments to solve non-convex mixed-integer nonlinear optimization problems.

[0097] In determining the control scenario and agent type, the scenario focuses on a flexible manufacturing workshop, which needs to cope with complex environments such as narrow passages, dense material areas, and alternating dynamic and static obstacles. The core objectives are high-precision formation control, obstacle avoidance, and low communication overhead. The agents are divided into a leader AGV and a follower AGV, which collaborate to complete coordinated material handling. The division of roles determines the differences in subsequent decision-making logic. The core elements of the model are defined (global state space and action space): The global state space integrates five types of key information, including the AGV's own pose, speed, formation parameters, upper bound of the communication cycle, and a dynamic obstacle avoidance risk map, forming ISCC global observation data to solve the problem of incomplete information from single local observations. The action space includes continuous velocity vectors (for acceleration, deceleration, and steering) and discrete communication cycles (for dynamically adjusting the communication frequency), covering both motion control and communication optimization dimensions to adapt to hybrid decision-making requirements.

[0098] Furthermore, a state transition function is constructed. This function takes the current global state and the agent's actions as input and outputs the probability distribution of the next state. Its core function is to quantify the impact of actions on state changes, such as the probability of AGV pose change after adjusting speed, or the probability of network state fluctuation after modifying the communication cycle, providing a probabilistic basis for decision-making.

[0099] Furthermore, differentiated reward functions are designed, including a reward function for the leading AGV and a reward function for the following AGV. The reward for the leading AGV consists of obstacle avoidance rewards (avoiding collision risks), navigation rewards (reducing travel time), and interaction rewards (optimizing communication overhead), guiding it to focus on path planning and efficiency improvement. The reward for the following AGV consists of obstacle avoidance rewards (ensuring safety), formation rewards (minimizing synchronization errors), and interaction rewards (controlling communication costs), with the core objective of maintaining formation accuracy. Differentiated rewards avoid gradient conflicts between heterogeneous agents while simultaneously considering multiple objectives such as safety, efficiency, and formation stability.

[0100] Furthermore, a CTDE framework is constructed and the Actor-Critic network is split. A centralized training and distributed execution (CTDE) framework is adopted. During the training phase, global information from the ISCC is centrally utilized to optimize the model. During the execution phase, each AGV makes autonomous decisions, balancing collaborative accuracy and real-time performance. Network splitting divides the Actor network into dedicated networks for the leading AGV and those for the following AGV, each learning its corresponding strategy. The Critic network is shared and responsible for evaluating the value of actions based on global information. This avoids gradient direction conflicts between different agents, accelerates training convergence, and improves the algorithm's generalization ability.

[0101] Finally, based on the differentiated reward function, the communication cycle allocation mechanism, and the CTDE framework, a multi-agent formation control algorithm based on ISCC is determined. The differentiated reward function, the dynamic communication cycle allocation mechanism, and the CTDE framework are integrated to form a complete POMDP decision model. This transforms the discrete + continuous decision problem of multi-agent formation control into a quantifiable and learnable probabilistic decision problem, adapting to sensory control constraints.

[0102] In one specific embodiment, the partially observable Markov process is as follows: Since the formation control optimization problem has multiple local optima and involves discrete and continuous decision variables, to handle this non-convex mixed-integer nonlinear optimization problem, this application formulates the multi-agent formation control problem as a partially observable Markov Decision Process (POMDP) ​​under communication, perception, and control constraints. This application proposes a MAPPO formation control strategy based on ISCC global information. Unlike traditional MAPPO, this application fully utilizes communication and perception system information, such as risk maps and interaction cycles, as ISCC state observations to improve the training efficiency of multi-agent formation control. Then, by learning the optimal communication strategy in a complex dynamic environment, the agents dynamically adjust their communication cycles to ensure synchronization accuracy and minimize communication overhead caused by additional ISCC global states. In ISCC-MAPPO, this application utilizes the high-dimensional state information of the ISCC system and the differential collaborative learning between heterogeneous leader and follower agents. Specifically, to achieve cooperation among heterogeneous agents, this application introduces a centralized training and distributed execution (CTDE) framework for training. Furthermore, unlike common shared networks, in the CTDE framework, this application splits the actor-critic network into two independent agent networks, separately computing the decisions of the leader and follower agents, thus avoiding gradient conflicts and improving training efficiency. Multiple rewards, such as obstacle avoidance, formation control, navigation, and interaction overhead, are employed to handle both continuous and discrete action spaces, thereby improving the performance of the ISCC-MAPPO algorithm.

[0103] Specifically, in addition to receiving... Partial observation data of itself Such as the AGV's own pose information driving speed In formation queue parameters It also supplements formation pose and global risk map data from multi-source sensors in the sensor network. Upper bound of the communication cycle of AGVs in a communication network The remaining observation data were merged into complete ISCC global state observation data. Send to This approach provides more comprehensive state information, which helps improve the efficiency and accuracy of model training and supports multi-agent formation control systems in handling tasks such as high-precision formation control in complex dynamic environments. Furthermore, to balance the additional communication overhead caused by interacting with ISCC global information, agents utilize a dynamic period allocation mechanism. Based on ISCC state observations, they can dynamically adjust their own communication interaction period, ensuring high-precision synchronization in formation control while significantly reducing communication overhead and closed-loop latency.

[0104] First, POMDP can be used The diagram illustrates the concept of collaborative transport in industrial environments. Here, N represents the number of agents, S is the global state, A is an action chosen by each agent, P is the state transition function, r is the global joint reward after the agent selects an action using the global state during the global training phase, O is the local environmental observation obtained by each agent, and Z is the observation function. Each AGV agent can only see its surrounding environment. The AGV agent needs to perceive information about its surrounding environment to update its confidence in its current state for future action selection decisions. Furthermore, collaborative transport in industrial environments is mainly accomplished through cooperation between two types of agents: Leader agents and Follower agents. The two types of agents have the same state and action spaces but different reward mechanisms. Therefore, different Actor networks need to be designed for the Leader and Follower to learn their respective policies, avoiding gradient direction conflicts between different types of agents, accelerating convergence, and improving the algorithm's dynamic adaptability and generalization ability. The Critic network can be shared so that each learns the policy best suited to its role.

[0105] 1) State Space: The state space is the foundation of a multi-agent formation control system. By accurately describing the local states of the AGV, it effectively reflects its real-time environmental and self-state information in complex environments. The state space of an agent should include various key local information. The state space consists of five elements: AGV's own pose information... driving speed In formation queue parameters Upper bound of communication period Global risk map data The local observation state of agent n and global information The definition is as follows:

[0106] ;

[0107] ;

[0108] 2) Action Space The motion space design should ensure that AGVs can learn and optimize their behavior strategies in constantly changing network topology and travel environments. In narrow corridors or dense material areas of a workshop, and in complex environments with alternating dynamic and static obstacles, AGVs need to continuously adjust their motion states to minimize travel time and achieve formation control optimization goals while meeting sensor-based control constraints. Therefore, the motion space design should consider the diverse situations that AGVs may encounter during task execution, enabling the agent to flexibly and quickly respond to environmental changes. The motion space of a multi-AGV system includes a continuous velocity vector and a discrete communication cycle.

[0109] Action space It can be represented as: This action space encompasses the main dimensions of requesting control commands and executing control actions during multi-AGV collaborative formation. By adjusting these parameters in the action space, multi-AGV agents can achieve various behavioral choices, including acceleration, deceleration, turning, and adjusting communication interaction cycles to adapt to dynamic network topology and obstacle avoidance risk areas. Joint Action Space The definition is as follows: ;

[0110] 3) Reward Function: To avoid gradient conflicts among heterogeneous agents during ISCC-MAPPO training, this application decouples the actor-critic network into two independent agent networks, which learn and train cooperative transport decisions separately. Furthermore, considering the need for AGV systems in flexible manufacturing workshops to balance obstacle avoidance, smooth driving, and precise synchronization, this application differentiates the reward functions for different types of agents. The Leader's reward consists of obstacle avoidance, navigation, and interaction rewards, aiming to guide it to the target point while avoiding obstacles and optimizing interaction overhead. The Follower's reward consists of obstacle avoidance, formation, and interaction rewards, aiming to maintain formation with the Leader while avoiding obstacles and optimizing interaction overhead.

[0111] a) Obstacle Avoidance Bonus: Considering the complex environment of narrow corridors or dense material areas in flexible manufacturing workshops, and the interplay of dynamic and static obstacles, the risk of collisions increases due to delayed closed-loop control commands and untimely obstacle avoidance braking. Based on this, according to the risk value... Set obstacle avoidance penalties, especially when the risk value is high. A larger negative reward is given when the obstacle is below a safe threshold, encouraging the AGV to maintain a safe distance from dynamically moving obstacles and plan routes in low-risk areas to avoid decreased transportation efficiency due to collisions or congestion. This reward function design helps to improve the global average travel speed and reduce travel time while maintaining driving safety. Obstacle Avoidance Reward Function It can be represented as:

[0112] ;

[0113] in, For obstacle avoidance sensitivity coefficient, Indicates confidence level The conditional value of risk (VoV) reflects the expected probability of a collision in the worst-case scenario. The Actor-Critic network in MAPPO can learn the underlying patterns of speed fluctuations under different risk values, thereby finding the adjustment range for smooth speed under different risk values ​​and suppressing potential hazards such as material compression and frictional deformation.

[0114] b) Formation Reward: Unlike the dynamically changing formations in other multi-agent formation control systems, the AGV formation in the flexible workshop is a static, strictly geometric formation. AGVs need to maintain high-precision synchronization with other AGVs in the formation to prevent material drop and damage. Therefore, formation control prioritizes maintaining the AGV formation and minimizing synchronization errors. When the synchronization error approaches zero, the positive incentive of the reward function is stronger. Once the synchronization error exceeds the maximum allowable positional deviation threshold of the formation, the penalty increases. Formation Reward It is expressed as follows:

[0115] ;

[0116] in, The formation sensitivity coefficient represents the degree to which formation rewards affect real-time rewards. This is the synchronization error threshold. The smoothness of speed changes is analyzed based on the variation in the synchronization error value to characterize the risks of material center-of-gravity shift, such as compression, friction, deformation, and even drop damage. Furthermore, when the AGV is the leader, formation error needs to be considered to adjust its own speed.

[0117] c) Navigation Reward: In flexible manufacturing workshops, travel time—the time it takes for an AGV to travel from its origin to its destination—is crucial to production efficiency. High communication latency and the limited perception range within certain areas significantly restrict the current AGV speeds within the workshop, increasing travel time. Therefore, when designing the reward function, while ensuring synchronization error and obstacle avoidance risks, the function encourages AGVs to increase their travel speed to reduce travel time and improve control efficiency. Navigation Reward Function The expression is as follows: ;in, The navigation sensitivity coefficient represents the degree to which navigation rewards affect immediate rewards. This represents the total distance traveled by the formation towards its destination at the current moment. This represents the total distance traveled by the AGV towards its destination in the previous moment. Navigation rewards only apply to the leader agent; when the AGV is a follower, there are no AGV formation rewards.

[0118] d) Interactive Rewards: Considering the limited network resources, the high-frequency uploading of sensing data and receiving of control decisions by a large number of intelligent machines within a unit of time can easily cause network congestion, making it difficult to meet the low-latency formation control requirements of large-scale AGVs. Therefore, based on previous ISCC research, an upper bound on the communication period is introduced. That is, under the stable condition of the control system, the maximum duration during which AGVs can maintain a safe distance and ensure the stability of the control system without sending sensing data to and receiving sensing and control data from the server. Unlike traditional control schemes with fixed communication interaction cycles, the reward function provides rewards to AGVs based on the dynamically adjusted communication interaction cycle in subsequent work of this application. While ensuring that pose error and synchronization error are within constraint thresholds, AGVs operating in open, low-risk areas are encouraged to adjust their communication cycle. This reduces the number of interactions with base stations, lowers the probability of data retransmission and communication latency, and reduces communication overhead. Simultaneously, it balances and improves the communication quality and environmental response speed of AGVs operating in high-risk areas. (Interactive reward function) The expression is as follows:

[0119] ;

[0120] in, The interaction coefficient represents the degree to which the reward in each communication cycle affects the immediate reward. It is the dynamic adaptive communication cycle of the AGV at time t. This is the maximum tolerable communication cycle for an AGV based on a safe distance, under stable control conditions derived from previous work. Furthermore, the communication cycle... Typically triggered and dynamically changed in open, low-density areas, other areas This represents the reference cycle time for AGVs performing cooperative formation tasks in the 3GPP standard. Based on this reward function, the Actor-Critic network can also learn the impact of communication cycles on formation synchronization errors under different risk value regions. This allows them to find ways to ensure controllable formation synchronization errors while not only saving significant communication overhead and enhancing communication network performance, but also releasing the controllability of AGVs, such as expanding the number of controllable formations and improving AGV environmental response and average travel speed.

[0121] In summary, the joint reward function of multiple agents in the formation at time t Represented as:

[0122] .

[0123] In summary, this application designs a MAPPO formation control system based on ISCC global information, which implements multi-agent formation control in two phases: training and execution. In the training phase, the agents first... At time t, the local observations are uploaded via the uplink. It receives ISCC global information observations from the base station via the downlink. Based on this information, the internal Actor network outputs actions in the continuous action space. The edge server generates a shared observation value based on the ISCC global information observations received from all current agents. The feedback is sent to the Critic network, and the estimated value of agent n in the current state is returned. Next, by... The experience pool of the agent is stored and randomly sampled to update the loss function parameters of the Actor network and Critic network. During the execution phase, each agent's Actor network updates the current ISCC global information observations received from interacting with the base station. It directly outputs the motion control strategy and a new communication interaction cycle. Furthermore, within each communication interaction cycle, each Agent supplements the ISCC global observation content based on its own local observations. Directly output motion control strategy and communication interaction cycle .

[0124] Specifically, based on the differentiated reward function, communication cycle allocation mechanism, and CTDE framework, a multi-agent formation control algorithm based on ISCC is determined. During the training phase, the Actor network continuously calculates the joint advantage function based on the global information of the agents in the Critic network in order to make effective action decisions. Update using the pruned loss function. Perform gradient ascent optimization and update network parameters. Its expression is as follows:

[0125] ;in, The ratio of the probabilities of the new strategy to the old strategy for the Actor represents the proportion coefficient of importance sampling. It's a clipping function used to limit... The range of changes should be adjusted to prevent the update step size from being too large. Indicates the cropping range. This means that the generalized dominance estimation method calculates the dominance function at each time step, which is used to measure the dominance in a certain state. Select a specific action The feasibility of this.

[0126] Joint advantage function It can be represented as: ;in, For state value function value, This represents the discount factor. Furthermore, the Actor network parameters for the leader and follower agents are dynamically updated based on the AGV agent type, while the Critic is typically shared, inputting global observations to estimate the value function, thus requiring no differentiation.

[0127] In the Critic network part, the Critic network mainly minimizes the loss function. Perform gradient descent optimization and update network parameters. The expression is: ;in, This indicates that the Critic network is related to the state. Value function estimation, It is a cumulative return, that is .

[0128] Optionally, based on the ISCC multi-agent formation control algorithm, networked formation control is performed on the AGV formation, including:

[0129] Initialize the global state reliability distribution of the AGV formation;

[0130] At each time step, each AGV integrates its own local observations with the ISCC global observations to form a complete observation, and updates the confidence distribution using Bayes' formula;

[0131] Based on the updated confidence distribution, and according to the state transition function and differentiated reward function in the ISCC multi-agent formation control algorithm, the motion control action and communication cycle decision are output.

[0132] In this embodiment, this step achieves precise formation control in dynamic environments by updating the confidence distribution and calling the model, balancing decision-making accuracy and real-time performance. During the initialization of the global state confidence distribution, the AGV lacks environmental observation information in the initial stage. Setting the probability distribution (i.e., confidence distribution) of the global state based on prior knowledge provides a foundation for subsequent initial decisions, addressing the problem of blind decision-making caused by insufficient initial observations.

[0133] At each time step, the AGV first collects its own local observations (pose, speed, formation parameters), and then receives ISCC global observations (global risk map, upper bound of communication cycle) from the edge server, forming complete observation data. The confidence distribution is updated using Bayes' theorem: combining new observations with prior confidence, the posterior probability distribution is calculated to quantify the uncertainty of the current global state. This adapts to the "partially observable" characteristic of POMDP, reducing state estimation errors and improving decision reliability through Bayesian inference. Based on the updated confidence distribution, the state transition function (predicting the impact of actions on the next state) and the differentiated reward function (evaluating the value of actions) in the ISCC-based multi-agent formation control algorithm are invoked. Two types of decisions are output: motion control actions (speed, steering adjustment) and communication cycle decisions (dynamically adjusting interaction frequency). This ensures that the decisions meet formation accuracy requirements while optimizing communication overhead based on environmental risks (e.g., increasing communication frequency in high-risk areas and decreasing frequency in low-risk areas).

[0134] Furthermore, the method also includes:

[0135] Based on the ISCC's multi-agent formation control algorithm, the lead AGV generates navigation guidance action decisions based on navigation state elements;

[0136] According to the ISCC multi-agent formation control algorithm, the following AGV generates formation-guided action decisions based on formation state elements; wherein, after receiving the action decisions, the edge server connected to the AGV formation sends out ISCC global observation supplementary data to update the AGV's state reliability distribution.

[0137] In this embodiment, this step achieves leader-follower coordination through role-differentiated decision-making and global information feedback, ensuring formation consistency and navigation efficiency. The navigation guidance decision of the leader AGV is based on navigation state elements (such as destination travel progress, global risk map, and upper bound of communication cycle) in the ISCC-based multi-agent formation control algorithm. The decision objective is to generate the optimal navigation path and speed strategy, shorten travel time while avoiding obstacles, and provide a stable formation reference for the following AGVs. Based on the navigation reward function, the speed is appropriately increased in low-risk areas, and the path is adjusted and a safe speed is maintained in obstacle-dense areas.

[0138] The decision-making basis for AGV formation guidance is the formation state elements (such as synchronization error with the lead AGV, formation parameters, and local obstacle avoidance information) in the ISCC-based multi-agent formation control algorithm. The decision objective is to generate speed and steering adjustment strategies, minimize the synchronization error with the lead AGV, maintain a strict geometric formation, and prevent material from falling. For example, when the synchronization error approaches a threshold, the speed is adjusted to reduce the gap; upon receiving obstacle avoidance observations, the path is fine-tuned while maintaining formation.

[0139] After receiving the action decisions from all AGVs, the edge server integrates multi-source sensor data to generate the latest ISCC global observation supplementary data (such as an updated global risk map and the overall pose of the formation). This supplementary data is then distributed to each AGV to update its global state confidence distribution, providing a precise basis for decisions at the next time step. This forms a closed loop of decision-making, observation, and updating, ensuring the dynamic adaptability of formation control and addressing the impact of environmental changes and communication delays.

[0140] This application proposes a closed-loop control method for multi-agent intelligent machines that integrates communication, sensing, and control. This method provides a safe, efficient, and networked control solution for the strictly consistent movement of large-scale multi-AGVs in dynamic and complex environments. It also offers a multi-agent formation control solution that ensures optimal control performance while significantly reducing communication overhead. By integrating sensing, communication, and control functions, and jointly designing a networked formation control system for multi-agent intelligent machines in flexible manufacturing workshops, this method avoids sacrificing movement speed to reduce the risk of material compression and deformation during emergency braking and startup. It balances synchronization accuracy and communication overhead, providing a safe, efficient, and scalable formation control paradigm for industrial collaborative material handling scenarios. Specific solutions involved include:

[0141] 1) Dynamic obstacle avoidance risk map based on CVaR. The perception methods at this stage are not limited to camera perception; any perception method that can obtain the AGV's position can be used as an alternative.

[0142] 2) Based on the dynamic communication cycle allocation mechanism of networked control state, the derivation method for solving the upper bound of the communication cycle in this stage is not limited to the perturbation method. Any method that derives the upper bound of the communication cycle based on the stability condition of the multi-AGV trajectory optimization equation can be used as an alternative.

[0143] 3) Partially observable Markov processes. The modeling methods at this stage are not limited to the standard POMDP definition. Any multi-agent formation control algorithm that can characterize the state transition and policy learning of multiple agents under incomplete information conditions can be used as an alternative.

[0144] 4) MAPPO formation control based on ISCC global information. The reinforcement learning solution method at this stage is not limited to the MAPPO algorithm. Any multi-agent reinforcement learning method that can integrate communication, perception and control global information within a centralized training-distributed execution (CTDE) framework can be used as an alternative.

[0145] Reference Figure 4 The diagram illustrates the variation of formation synchronization accuracy for different algorithms within a dynamic and complex environment. Figure 4The performance of traditional MAPPO control strategy, traditional MADDPG control strategy, traditional virtual structure-based formation control strategy, leader-follower-based formation control strategy, and the multi-agent formation control system designed in this application were compared. As shown in the figure, the strategy proposed in this application can maintain a stable synchronization accuracy of 10mm, significantly outperforming the learning-based traditional MAPPO and MADDPG strategies. This is because this application designs the formation control strategy from the perspective of communication and perception-assisted enhanced control, fully utilizing global perception information to construct a CVaR map to avoid potential collision risks, significantly reducing the probability of emergency braking, and reserving sufficient time for smooth speed changes and formation synchronization. Furthermore, the multi-agent dynamically adjusts the interaction cycle between the AGV and the base station based on ISCC state observations, improving communication resource utilization and reducing communication latency. This means that the multi-agent acquires dynamic environmental changes with extremely low latency and evaluates changes in formation structure in real time, thereby achieving more accurate synchronization decisions and improving synchronization accuracy and control efficiency. In addition, from... Figure 4 As can be seen, the leader-follower strategy can maintain high-precision formation synchronization in the first 25 seconds. However, as time increases, the algorithm struggles to effectively guarantee a strict geometric formation. This is because model-based formation control methods rely on pre-set rules for control decisions, maintaining system stability in environments conforming to these rules. When the environment is unpredictable or dynamically changing, model inaccuracies and high communication latency affect control performance, failing to guarantee the effectiveness of all decisions in complex and unknown environments. Therefore, this also verifies that the proposed algorithm can fully characterize the coupling relationship between communication, perception, and control, effectively expand the number of controllable agents, reduce communication overhead and latency, and further improve control performance such as navigation and obstacle avoidance.

[0146] Reference Figure 5 The diagram shows the trajectory change pattern of a multi-agent formation in a complex environment with alternating dynamic and static obstacles. Figure 5 The top part of the diagram shows the movement trajectories of the multiple agents. Figure 5 The middle section of the diagram illustrates the variation of linear velocity among multiple agents. Figure 5 The lower part of the diagram illustrates the variation pattern of synchronization accuracy in multi-agent formations.

[0147] This application simulates the process of 28 AGV nodes performing a collaborative transport task in an unknown dynamic environment. The movement trajectories and motion state changes of the four AGVs in the first diamond formation are shown below. Figure 4 As shown. Figure 5 In the top section, the X and Y axes represent the workshop's relative two-dimensional coordinates. From... Figure 5The top part shows that after receiving the coordinated transport request at the initial moment, Formation 1 began to reduce its speed to maintain formation and search for the shortest path to its destination. After 55 seconds, from... Figure 5 The speed fluctuation curve in the middle section shows that the speed of Formation 1 decreases sharply. This is because, based on the acquired global CVaR risk map information, this formation anticipates passing through point (50, 70) on its upcoming path. Since an unknown static obstacle exists near this point within the safe distance warning range of Formation 1's planned path, other AGV formations near this point are considered moving obstacles, posing a high potential risk of congestion and collision. Therefore, the leader within Formation 1, unable to determine whether a congestion collision will occur, begins a slow and smooth deceleration in advance to avoid the risk of the transported objects shifting their center of gravity and being damaged due to drastic speed fluctuations. Other follower agents within the formation also synchronously decelerate to maintain strict geometric synchronization.

[0148] After 80 seconds, as Formation 1 approached the obstacle area, it gradually accelerated to its maximum speed. Then, at 140 seconds, Formation 1 passed through (50, 70) again. Having learned optimal control decisions in similar environments, Formation 1 determined this to be a safe obstacle area. At this point, the formation decided to pass through without slowing down, significantly increasing its adaptability to unknown and complex dynamic environments and improving its driving efficiency.

[0149] In addition, from Figure 5 The bottom part shows that when the AGV formation first reaches its maximum speed of 5 m / s around 25 seconds, the formation synchronization accuracy of some AGV formations is close to the given safety threshold of 30 mm. After 120 seconds, when it reaches its maximum speed of 5 m / s again, the formation synchronization accuracy of all formations stabilizes at around 10 mm. This is because, on one hand, the proposed algorithm uses ISCC global information to construct a CvaR risk map, quantifying the severity of potential congestion and collisions between AGVs and dynamic obstacles during closed-loop formation control. This allows sufficient time for smooth speed changes and formation synchronization, eliminating the need for agents to sacrifice movement speed to reduce the risk of material compression and deformation during emergency braking and startup. Simultaneously, through a dynamic communication cycle allocation mechanism based on networked control state, high-precision synchronization of formation control is ensured while reducing communication overhead, avoiding synchronization errors caused by lagging control decisions and perception information leading to large pose errors. On the other hand, the proposed algorithm, through high-dimensional ISCC state information and differentiated collaborative learning among heterogeneous agents, can learn the optimal control strategy from historical environmental interactions, significantly improving formation synchronization accuracy and driving efficiency.

[0150] On the other hand, this application also provides a performance comparison of multi-agent formation control. Table 1 shows a comparison of parameters such as closed-loop delay and maximum number of AGVs supported between traditional solutions such as MAPPO and MADDPG and the design scheme of this application.

[0151] Table 1: Comparison of Multi-Agent Formation Control Performance

[0152]

[0153] As shown in Table 1, when multiple agents perform a cooperative transport task along a 100m long path, the average travel time is 33.03s, which is at least 16.8% lower than the baseline algorithm. This is because, unlike the traditional MAPPO and MADDPG baseline algorithms that utilize low-dimensional motion state information such as position and velocity, the proposed algorithm focuses more on the interaction between communication, perception, and control processes. This means that the formation can fully utilize the rich changes in ISCC global information to learn the optimal strategy under different environments. Agents do not need to sacrifice movement speed to reduce the risk of material compression and deformation during emergency braking and startup, greatly improving the formation's travel speed and thus reducing travel time.

[0154] In Table 1, cycle time represents the communication interval between the agent and the base station. A larger value indicates a smaller communication traffic per unit time, which can be approximately understood as lower network communication overhead. As shown in Table 1, this application, by dynamically adjusting the cycle time, reduces communication overhead by at least 15.5% compared to the baseline algorithm, alleviating communication pressure. During networked formation control, closed-loop latency is reduced by at least 52.7%, significantly improving communication performance. This is because the proposed algorithm's multi-agent system can learn the optimal control strategy based on rich ISCC global information, and also learn the optimal communication strategy that meets the current control requirements. Through the designed dynamic communication cycle allocation mechanism for networked control states, the AGV can dynamically output appropriate communication interaction cycles within the action space based on ISCC state observations, such as motion state and congestion collision risk values ​​in the driving area, reducing unnecessary information interaction and ensuring high-precision synchronization of formation control while significantly reducing communication overhead and closed-loop latency. Traditional MAPPO and MADDPG baseline algorithms neglect the requirements of communication networks for control performance under different complex environments, and the existing fixed-cycle communication strategies are not suitable for complex formation tasks in dynamic environments. Specifically, unlike the complex obstacle environment and dense dynamic obstacles in the assembly area, multi-agent vehicles do not need frequent interaction with the base station for strict obstacle avoidance and navigation when navigating in open, non-assembly areas, thus having lower communication requirements. However, the baseline algorithm's communication strategy still allocates the same communication resources to the multi-agent vehicles in this state, resulting in high-frequency interaction with redundant data, which leads to a significant waste of communication resources, increased communication latency, and delayed control decisions and perception information, causing a decrease in control efficiency and safe obstacle avoidance capabilities.

[0155] Furthermore, the proposed solution outperforms benchmark algorithms in key metrics such as formation synchronization accuracy and AGV speed, and supports larger-scale AGV formations, demonstrating superior scalability. These advantages stem from the multi-dimensional optimization of the high-precision formation control process through an integrated sensing and control design. This allows multiple agents to acquire rich ISCC global information with lower communication overhead, enabling them to quickly and efficiently learn optimal control and communication strategies in dynamic and complex environments. This verifies that the proposed algorithm can provide a safe, efficient, and scalable formation control paradigm for industrial collaborative material handling scenarios.

[0156] The various methods of the embodiments of this application have been described above. Apparatus for implementing the above methods will now be provided.

[0157] Please refer to Figure 6 This application also provides a networked formation control device integrating communication, sensing, and control, comprising:

[0158] The first construction module 61 is used to construct a dynamic obstacle avoidance risk map based on the conditional value at risk model;

[0159] The first determining module 62 is used to determine the upper bound of the communication cycle based on the stability condition of the trajectory optimization equation of the automated guided vehicle (AGV) formation, and to determine the communication cycle allocation mechanism of the AGV formation based on the upper bound of the communication cycle and the dynamic obstacle avoidance risk map.

[0160] The second determining module 63 is used to determine the ISCC-based multi-agent formation control algorithm based on the dynamic obstacle avoidance risk map and the communication cycle allocation mechanism, combined with the ISCC global information of the integrated communication perception and control.

[0161] The first processing module 64 is used to perform networked formation control on the AGV formation according to the ISCC multi-agent formation control algorithm.

[0162] Optionally, the first building module 61 described above includes:

[0163] The first acquisition unit is used to acquire the interaction parameters between the current AGV formation and dynamic obstacles; the dynamic obstacles are used to represent AGVs in the neighboring formations of the current AGV formation.

[0164] The first processing unit is configured to calculate the relative speed and safe braking distance between the AGV and the dynamic obstacle based on the interaction parameters, under the boundary condition that the AGV in the current AGV formation and the dynamic obstacle simultaneously decelerate at maximum speed and collide when they come to a stop.

[0165] The first determining unit is used to determine the congestion collision probability expression based on the relative speed and the safe braking distance;

[0166] The second determining unit is used to determine the objective function of the dynamic obstacle avoidance risk map based on the congestion collision probability expression.

[0167] The first construction unit is used to construct a dynamic obstacle avoidance risk map based on N sets of collision probability samples and the objective function.

[0168] Optionally, the first determining module 62 described above includes:

[0169] The second building unit is used to construct dynamic trajectory optimization equations based on the motion process of AGV formations traveling sequentially along the circular production line and the preset dynamic model.

[0170] The second processing unit is used to calculate the safe distance and optimal speed parameters under the business cycle based on the stability conditions of the trajectory optimization equation, the AGV pose information and the dynamic obstacle avoidance risk map, and to determine the upper bound of the communication cycle for maintaining the stability of the closed-loop interaction of the formation control system. The stability conditions are used to determine that the following AGV can track the trajectory of the leading AGV, so that the position error and speed error within the formation converge to zero or are limited to a safe threshold.

[0171] Optionally, the first determining module 62 described above includes:

[0172] The third processing unit is used to determine the current area as a low-risk area when it is a non-assembly area, and as a high-risk area when it is an assembly area, based on the dynamic obstacle avoidance risk map.

[0173] The fourth processing unit is used to adaptively increase the communication cycle in the communication cycle when the low-risk area does not exceed the upper limit of the communication cycle, and to determine the communication cycle as an integer multiple of the AGV control business cycle through the cofunction; when the closed-loop control delay is less than the communication cycle, the AGV continues to execute the control decision of the previous cycle according to the updated motion state.

[0174] The fifth processing unit is used to adjust the communication cycle to a preset communication cycle in the case of the high-risk area, and to obtain a global risk map and high-precision pose information through high-frequency base station interaction.

[0175] Optionally, the second determining module 63 described above includes:

[0176] The third determining unit is used to determine the AGV formation control scenario and the type of intelligent agent; the type of intelligent agent includes leading AGV and following AGV.

[0177] The fourth determining unit is used to determine the global state space and action space; the global state space is the integrated ISCC global information of communication, perception and control determined based on the AGV pose, driving speed, formation parameters, the upper bound of the communication cycle and the dynamic obstacle avoidance risk map.

[0178] The sixth processing unit is used to construct a state transition function based on the global state space and the action space; the state transition function takes the current state and the current action as input and the probability distribution of the next state as output;

[0179] The third construction unit is used to construct the intelligent agent type and construct a differentiated reward function; the differentiated reward function includes a leading AGV reward function and a following AGV reward function; the leading AGV reward function includes obstacle avoidance, navigation and corresponding interaction rewards for the leading AGV; the following AGV reward function includes obstacle avoidance, formation and corresponding interaction rewards for the following AGV.

[0180] The seventh processing unit is used to split the execution-reward Actor-Critic network corresponding to the multi-agent reinforcement learning algorithm into a leading AGV Actor network, a following AGV Actor network, and a shared Critic network, and to build a centralized training and decentralized execution CTDE framework.

[0181] The fourth construction unit is used to determine the ISCC-based multi-agent formation control algorithm based on the differentiated reward function, the communication cycle allocation mechanism, and the CTDE framework.

[0182] Optionally, the first processing module 64 described above includes:

[0183] The eighth processing unit is used to initialize the global state confidence distribution of the AGV formation;

[0184] The ninth processing unit is used at each time step. Each AGV integrates its own local observations with the ISCC global observations to form a complete observation, and updates the confidence distribution through Bayes' formula.

[0185] The tenth processing unit is used to output motion control actions and communication cycle decisions based on the updated confidence distribution and according to the state transition function and differentiated reward function in the ISCC multi-agent formation control algorithm.

[0186] Optionally, the apparatus of this application further includes:

[0187] The second processing module is used to generate navigation guidance action decisions for the AGV based on navigation status elements according to the ISCC multi-agent formation control algorithm.

[0188] The third processing module is used to generate formation-guided action decisions based on the formation state elements of the AGV and the multi-agent formation control algorithm of the ISCC. The edge server connected to the AGV formation receives the action decisions and sends out supplementary global observation data of ISCC to update the state reliability distribution of the AGV.

[0189] It should be noted that the device in this embodiment corresponds to the networked formation control method integrating communication, sensing, and control described above. The implementation methods in the above embodiments are all applicable to the embodiments of this device and can achieve the same technical effect. The device provided in this application embodiment can implement all the method steps implemented in the above method embodiments and can achieve the same technical effect. Therefore, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail here.

[0190] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described networked formation control method embodiment integrating communication, sensing, and control, and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0191] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described networked formation control method embodiment that integrates communication, sensing, and control, and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0192] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0193] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0194] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A networked platooning control method of communication and perception control integration, characterized by, include: Construct a dynamic obstacle avoidance risk map based on the conditional value-at-risk model; Based on the stability condition of the trajectory optimization equation of the AGV formation, the upper bound of the communication cycle is determined, and based on the upper bound of the communication cycle and the dynamic obstacle avoidance risk map, the communication cycle allocation mechanism of the AGV formation is determined. Based on the dynamic obstacle avoidance risk map and the communication cycle allocation mechanism, and combined with the ISCC global information of the integrated communication, perception and control system, a multi-agent formation control algorithm based on ISCC is determined. The AGV formation is networked and controlled according to the ISCC multi-agent formation control algorithm.

2. The method of claim 1, wherein, Constructing a dynamic obstacle avoidance risk map based on the conditional value-at-risk model, including: Collect interaction parameters between the current AGV formation and dynamic obstacles; the dynamic obstacles are used to represent AGVs in neighboring formations of the current AGV formation. Based on the interaction parameters, under the boundary condition that the AGV in the current AGV formation and the dynamic obstacle simultaneously decelerate at maximum speed and collide just when they come to a stop, the relative speed and safe braking distance between the AGV and the dynamic obstacle are calculated. Based on the relative speed and the safe braking distance, determine the congestion collision probability expression; Based on the congestion collision probability expression, determine the objective function of the dynamic obstacle avoidance risk map; A dynamic obstacle avoidance risk map is constructed based on N sets of collision probability samples and the objective function.

3. The method of claim 1, wherein, Based on the stability conditions of the trajectory optimization equation for AGV formations, the upper bound of the communication period is determined, including: Based on the motion process of AGV formations moving sequentially along a circular production line and a pre-set dynamic model, dynamic trajectory optimization equations are constructed. Based on the stability conditions of the trajectory optimization equation, the AGV pose information, and the dynamic obstacle avoidance risk map, the safe distance and optimal speed parameters under the business cycle are calculated, and the upper bound of the communication cycle for maintaining the stability of the formation control system is determined. The stability conditions are used to determine that the following AGV can track the trajectory of the leading AGV, so that the position error and speed error within the formation converge to zero or are limited to a safe threshold.

4. The method of claim 1, wherein, Based on the upper bound of the communication cycle and the dynamic obstacle avoidance risk map, the communication cycle allocation mechanism for the AGV formation is determined, including: According to the dynamic obstacle avoidance risk map, if the current area is a non-assembly area, it is determined to be a low-risk area; if the current area is an assembly area, it is determined to be a high-risk area. In the low-risk area, and within the upper limit of the communication cycle, the communication cycle is adaptively increased, and the communication cycle is determined to be an integer multiple of the AGV control business cycle by using the cofunction; when the closed-loop control delay is less than the communication cycle, the AGV continues to execute the control decision of the previous cycle according to the updated motion state. In the case of the high-risk area, the communication cycle is adjusted to a preset communication cycle, and a global risk map and high-precision pose information are obtained through high-frequency base station interaction.

5. The method of claim 1, wherein, Determine an ISCC-based multi-agent formation control algorithm based on the dynamic obstacle avoidance risk map and the communication cycle allocation mechanism, in combination with ISCC global information of integrated communication and perception control, including: Determine the AGV formation control scenario and the agent type; the agent type includes a leading AGV and a following AGV; Determine the global state space and the action space; the global state space is determined according to the AGV pose, the travel speed, the formation parameter, the upper bound of the communication cycle and the dynamic obstacle avoidance risk map, and the ISCC global information of integrated communication and perception control; According to the global state space and the action space, a state transition function is constructed; the state transition function takes the current state and the current action as input, and outputs the next state probability distribution; Construct the agent type, and construct a differentiated reward function; the differentiated reward function includes a leading AGV reward function and a following AGV reward function; the leading AGV reward function includes leading AGV obstacle avoidance, navigation and corresponding interaction rewards; the following AGV reward function includes following AGV obstacle avoidance, formation and corresponding interaction rewards; Split the execution-reward Actor-Critic network corresponding to the multi-agent reinforcement learning algorithm into a leading AGV Actor network, a following AGV Actor network and a shared Critic network, and construct a centralized training and decentralized execution (CTDE) framework; Determine an ISCC-based multi-agent formation control algorithm according to the differentiated reward function, the communication cycle allocation mechanism and the CTDE framework.

6. The method of claim 1, wherein, According to the ISCC multi-agent formation control algorithm, networked formation control is performed on the AGV formation, including: Initialize the global state credibility distribution of the AGV formation; At each time step, each AGV fuses its own local observation with the ISCC global observation to form a complete observation, and updates the credibility distribution through the Bayesian formula; Based on the updated credibility distribution, the state transition function and the differentiated reward function in the ISCC multi-agent formation control algorithm are used to output motion control actions and communication cycle decisions.

7. The method of claim 1, wherein, The method further includes: According to the ISCC multi-agent formation control algorithm, the leading AGV generates navigation guidance action decisions based on navigation state elements; According to the ISCC multi-agent formation control algorithm, the following AGV generates formation guidance action decisions based on formation state elements; wherein after the edge server connected to the AGV formation receives the action decisions, it issues ISCC global observation supplementary data to update the state credibility distribution of the AGV.

8. A networked platooning control apparatus of integrated communication and perception control, characterized by, Including: A first construction module for constructing a dynamic obstacle avoidance risk map based on a conditional risk value model; A first determination module for determining an upper bound of a communication cycle according to a stability condition of a trajectory optimization equation of an automated guided vehicle (AGV) formation, and determining a communication cycle allocation mechanism for the AGV formation according to the upper bound of the communication cycle and the dynamic obstacle avoidance risk map; A second determination module is configured to determine an ISCC-based multi-agent formation control algorithm based on the dynamic obstacle avoidance risk map, the communication cycle allocation mechanism, and integrated ISCC global information. A first processing module is configured to perform networked formation control on the AGV formation according to the ISCC-based multi-agent formation control algorithm.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 7.

10. A computer program product, characterised in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 7. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 7.