Motion coordination control system, method, terminal and medium based on machine learning and tsn deterministic communication
Patent Information
- Application Number
- CN202610947593.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-06-29
AI Technical Summary
[0008](1)现有回收设备通信架构在带宽和确定性方面存在不足:1553B总线1Mbps的带宽在多执行机构高频协同控制场景下严重不足,SpaceWire缺乏标准化的确定性保障机制,私有以太网缺乏互操作性生态
[0030] First, communication resource utilization is significantly improved. Through the joint optimization framework, the TSN communication resource configuration adaptively adjusts according to the control strategy, avoiding the resource waste caused by traditional static configuration. Simulation results show that, compared with the hierarchical scheme of fixed TSN configuration paired with a fixed AI controller, this joint scheme improves control performance by more than 15% and communication resource utilization by more than 30%.
Smart Images

Figure CN122450166B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of motion control technology, and in particular to motion cooperative control systems, methods, terminals and media based on machine learning and TSN deterministic communication. Background Technology
[0002] Regenerative braking in reusable transport and recovery devices is one of the most challenging technical aspects in the field of intelligent motion control. The entire recovery process requires millisecond-level coordination among multiple mechanisms, including continuous adjustment of power output, pneumatic or attitude control, and timing of landing / docking mechanism deployment and locking. This places extremely high demands on the deterministic latency assurance capabilities of the equipment's communication network and the real-time decision-making capabilities of the control system.
[0003] In recent years, Time-Sensitive Networking (TSN) and Artificial Intelligence (AI) control technologies have made some progress in the fields of device communication and data recycling control, respectively. However, these two technology systems have long adopted a layered and fragmented design approach and have not yet achieved deep integration and optimization. TSN is a collective term for a series of Ethernet enhancement standards developed by the IEEE 802.1 working group, including the time synchronization standard IEEE 802.1AS, the time-aware shaping standard IEEE 802.1Qbv, the frame preemption standard IEEE 802.1Qbu, and the frame duplication and elimination reliability standard IEEE 802.1CB. TSN provides bounded latency and low jitter deterministic transmission guarantees for critical data streams through time-triggered scheduling mechanisms, and has mature applications in industrial control and automotive fields.
[0004] In the field of precision motion equipment communication, existing patent CN122069238A divides the network into three categories based on the criticality and real-time requirements of data flow: periodic time-sensitive data flow, high-speed time-sensitive data flow, and best-effort data flow. It achieves transparent interconnection between the 1553B bus and the TSN switching network through an intelligent protocol conversion gateway, constructing a three-layer partitioned and dual-network converged avionics network system. This solution solves the problem of insufficient bandwidth in the traditional 1553B bus and achieves unified transmission of multiple service flows at the communication network architecture level. Furthermore, existing patent CN122070686A uses artificial intelligence to update the control law of the device's transmission control signals to compensate for transmission delays, supporting time-sensitive deterministic communication based on the TSN mechanism. This solution introduces AI into delay compensation in the TSN communication link, but its application scenario is 5G communication networks and does not involve the field of motion device recovery control.
[0005] In the field of AI-based recovery control, the industry has widely carried out research on AI control algorithms such as deep reinforcement learning landing / docking optimization and model prediction trajectory tracking. Compared with traditional PID control, these algorithms have better recovery accuracy and energy utilization in simulation environments. However, these AI control algorithms are generally based on a common assumption: that the communication link is idealized, sensor data can reach the controller instantly, and control commands can reach the actuator instantly.
[0006] In the field of safety control, existing patent CN122018436A discloses a "Multi-spindle Concurrent Machining Obstacle Avoidance System and Method Based on Digital Twin and CBF". This patent utilizes CBF to construct high-order safety constraints and generates safety correction control commands by solving quadratic programming. This method demonstrates the effectiveness of CBF in engineering safety control, but its application is limited to the field of CNC machining and is not combined with TSN communication or motion device recovery scenarios.
[0007] Therefore, the existing technology still has the following obvious shortcomings:
[0008] (1) Existing recycling equipment communication architectures are deficient in terms of bandwidth and determinism: the 1Mbps bandwidth of the 1553B bus is severely insufficient in high-frequency collaborative control scenarios involving multiple actuators; SpaceWire lacks a standardized deterministic guarantee mechanism; and private Ethernet lacks an interoperability ecosystem. Even the TSN communication scheme disclosed in CN122069238A has static or semi-static communication resource configurations, which cannot be adjusted in real time according to the dynamic changes in the recycling stage, leaving considerable room for improvement in communication resource utilization.
[0009] (2) Existing AI control schemes ignore the impact of communication non-idealization: Existing research on AI control for motion device retrieval assumes that sensor data can reach the controller instantaneously and control commands can reach the actuator instantaneously, without explicitly modeling the impact of communication delay, jitter, and packet loss on control performance. In actual systems, communication jitter and packet loss can directly lead to performance degradation or even instability of the AI controller. This problem is particularly prominent in scenarios where the control cycle at the retrieval end is compressed to 2 to 5 milliseconds.
[0010] (3) Layered decoupling of communication and control leads to suboptimal global performance: In traditional schemes, network engineers are responsible for designing the TSN communication network, and control engineers are responsible for designing the controller. The two subsystems are optimized independently. Since there is a strong coupling relationship between communication resource configuration and control strategy, the layered independent design cannot capture the interaction effect between the two, and the overall performance is inevitably inferior to the joint optimization scheme.
[0011] Based on the current state of technology, TSN communication and AI recycling control are currently designed in a hierarchical and independent manner. TSN communication focuses on the communication network architecture without considering the adaptation requirements of the control strategy, while AI recycling control assumes idealized communication without modeling actual communication constraints. Furthermore, no existing technology jointly solves for the deterministic TSN communication resource allocation and AI control strategy within the same optimization framework. Summary of the Invention
[0012] In view of the shortcomings of the prior art described above, the purpose of this application is to provide a solution to the following problems:
[0013] (1) How to provide end-to-end deterministic communication guarantee for high-frequency collaborative control of multiple actuators during the recovery process of the motion device, so that the allocation of communication resources can be adaptively adjusted with the dynamic changes of the recovery stage.
[0014] (2) How to enable the AI controller to explicitly perceive the actual capabilities of the communication link (including the upper limit of latency, jitter characteristics and packet loss rate) and maintain stable control performance under non-ideal communication conditions.
[0015] (3) How to solve the TSN communication resource configuration and AI control strategy parameters together under the same optimization framework, break through the performance bottleneck of traditional hierarchical independent design, and achieve global optimization of communication-control collaboration.
[0016] To achieve the above and other related objectives, a first aspect of this application provides a motion cooperative control system based on machine learning and TSN deterministic communication, comprising: a task planning layer, including a task planner; the task planner generates a full recovery task path based on planning basis information and outputs prior information of the recovery phase to the joint optimization layer at a preset period; the joint optimization layer includes a TSN-AI joint optimization engine; the TSN-AI joint optimization engine is connected to the task planner and includes a network situation awareness module and a joint optimization solver; the network situation awareness module is used to construct a real-time network situation representation; the joint optimization solver receives the network situation representation and the prior information of the recovery phase, sets a joint objective, and adopts a two-level decomposition. The strategy solves for the optimal GCL configuration template and optimal AI control parameters; the communication control execution layer includes a TSN deterministic communication subsystem, an AI recovery control subsystem, and a multi-actuator coordination subsystem; the TSN deterministic communication subsystem receives the optimal GCL configuration template and completes data transmission for the entire system based on clock synchronization, lossless switching of multi-condition gating lists, and dual-link hot backup transmission; the AI recovery control subsystem generates the optimal control quantity based on the optimal AI control parameters and real-time motion state estimation; the multi-actuator coordination subsystem receives the optimal control quantity and generates corresponding servo control commands; the physical perception layer includes sensor groups and actuator groups; both sensor groups and actuator groups transmit data through the TSN deterministic subsystem.
[0017] In some embodiments of the first aspect of this application, the task planner generates a full recovery task path based on planning basis information and outputs recovery stage prior information to the joint optimization layer at a preset period, including: the task planner generates a full recovery task path based on target location points, wind field predictions, and fuel remaining information; the generation of the full recovery task path includes calculating the reference motion trajectory of the motion device and defining the switching time nodes between each recovery sub-stage; the recovery stage prior information includes the current recovery sub-stage identifier and the task objective.
[0018] In some embodiments of the first aspect of this application, the network situation awareness module continuously collects actual end-to-end latency, network load status, and control performance indicators to construct a real-time network situation characterization.
[0019] In some embodiments of the first aspect of this application, the two-layer decomposition strategy adopted by the joint optimization solver includes: upper-layer recycling stage-level optimization and lower-layer frame-level real-time scheduling, wherein: the upper-layer recycling stage-level optimization process includes: running a Bayesian optimization algorithm at a preset period, and modeling the joint objective function using a Gaussian process as a surrogate model, and searching for the optimal GCL configuration template and optimal AI controller parameters by improving the expected acquisition function; the lower-layer frame-level real-time scheduling process includes: performing frame-level deterministic scheduling at a preset frequency based on the optimal GCL configuration template provided by the upper-layer recycling stage-level optimization, so as to dynamically adjust the guard band size and micro-time slot boundary without changing the GCL configuration template topology.
[0020] In some embodiments of the first aspect of this application, the step of running the Bayesian optimization algorithm at a preset period and modeling the joint objective function using a Gaussian process as a surrogate model, and searching for the optimal GCL configuration template and optimal AI controller parameters by improving the expected acquisition function, includes: compressing high-dimensional communication parameters and control parameters based on the prior information of the retrieval phase, combined with GCL parameterization and MPC parameterization; fitting the mapping relationship between parameters and optimization objectives based on network situation characterization, with control performance, resource efficiency, and security margin as joint optimization objectives, and using a Gaussian process as a surrogate model; and using the improved expected acquisition function as the search criterion, iteratively optimizing in the dimensionality-reduced parameter space using the Bayesian optimization algorithm to select the GCL configuration template and AI controller parameters that optimize the joint optimization objective.
[0021] In some embodiments of the first aspect of this application, the TSN deterministic communication subsystem includes a gated list manager, a redundant dual-network switching matrix module, and a time synchronization module; the gated list manager and the time synchronization module are respectively connected to the redundant dual-network switching matrix module, and the gated list manager receives the optimal GCL configuration template.
[0022] In some embodiments of the first aspect of this application, the AI recovery control subsystem includes an NN-MPC controller, a CBF safety filter, a safety RL policy network, and a switching decision module; the NN-MPC controller is connected to the CBF safety filter and the switching decision module, and is also connected to the joint optimization solver to obtain optimal AI control parameters; the safety RL policy network is also connected to the CBF safety filter and the switching decision module; wherein:
[0023] The NN-MPC controller takes motion state estimation data as input, learns the dynamic model residuals through a neural network, and solves for the optimal control quantity in the rolling time domain. The CBF safety filter is used to correct the optimal control quantity output by the NN-MPC controller. The safety RL policy network is used to automatically take over when the NN-MPC controller fails. The switching decision module switches between three control modes—NN-MPC dominant mode, parallel fusion mode, and safety RL independent mode—based on the neural network residual prediction error. In the parallel fusion mode, the NN-MPC controller and the safety RL policy network synchronously output control quantities and then weighted and fused them.
[0024] In some embodiments of the first aspect of this application, the switching strategy of the switching decision module includes: if the duration of the neural network residual prediction error exceeding a preset threshold does not exceed a preset alarm period, then the NN-MPC dominant mode is maintained; if the duration of the neural network residual prediction error exceeding the preset threshold continuously exceeds the preset alarm period, then the parallel fusion mode is switched, and when the NN-MPC controller cannot solve normally, the safe RL independent mode is switched.
[0025] To achieve the above and other related objectives, a second aspect of this application provides a motion cooperative control method based on machine learning and TSN deterministic communication, comprising: activating a task planner to generate a full recovery task path based on planning basis information, and outputting prior information of the recovery phase to the joint optimization layer at a preset period; activating the TSN-AI joint optimization engine to construct a real-time network situation representation, and based on the network situation representation and prior information of the recovery phase, setting a joint objective and using a two-layer decomposition strategy to solve for the optimal GCL configuration template and optimal AI control parameters; activating the TSN deterministic communication subsystem to receive the optimal GCL configuration template, and completing the data transmission of the entire system based on clock synchronization, lossless switching of multi-condition gating lists, and dual-link hot backup transmission; and activating the AI recovery control subsystem to generate an optimal control quantity based on the optimal AI control parameters and real-time motion state estimation; the multi-actuator cooperative subsystem receiving the optimal control quantity and generating corresponding servo control commands; activating the sensor group to collect raw sensing data of the motion state, and using the actuator group to perform corresponding operations.
[0026] To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the motion cooperative control method based on machine learning and TSN deterministic communication.
[0027] To achieve the above and other related objectives, a fourth aspect of this application provides a computer program product comprising computer program code that, when executed on a computer, enables the computer to implement the motion cooperative control method based on machine learning and TSN deterministic communication.
[0028] To achieve the above and other related objectives, a fifth aspect of this application provides an electronic terminal, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the motion cooperative control method based on machine learning and TSN deterministic communication.
[0029] As described above, the motion cooperative control system, method, terminal, and medium based on machine learning and TSN deterministic communication of this application have the following beneficial effects:
[0030] First, communication resource utilization is significantly improved. Through the joint optimization framework, the TSN communication resource configuration adaptively adjusts according to the control strategy, avoiding the resource waste caused by traditional static configuration. Simulation results show that, compared with the hierarchical scheme of fixed TSN configuration paired with a fixed AI controller, this joint scheme improves control performance by more than 15% and communication resource utilization by more than 30%.
[0031] Second, control robustness is significantly enhanced. The NN-MPC controller explicitly models communication constraints, and the control strategy parameters can be adapted to the actual latency and jitter characteristics of the communication link, maintaining stable control performance even under non-ideal communication conditions. The CBF safety filter provides formal safety guarantees, ensuring that the system state remains within a safe range even if the AI controller exhibits abnormal output.
[0032] Third, the overall performance achieves Pareto optimality. Traditional hierarchical schemes optimize communication and control modules independently, failing to effectively utilize their coupling relationship. This invention, based on a joint optimization framework, incorporates communication resource allocation and control strategy parameters into the same objective function, yielding a Pareto optimal solution in the communication-control joint space, with overall performance superior to various hierarchical independent optimization schemes.
[0033] Fourth, fault tolerance is significantly improved. Traditional solutions rely on conservative redundancy mechanisms to handle faults, while this invention uses dynamic reconfiguration to achieve efficient fault tolerance. The NN-MPC and security RL form a dual-mode redundancy architecture, which can complete controller switching within 5 milliseconds; the TSN dual-network redundancy can complete network switching within 2 milliseconds, and the switching response speed is significantly better than that of traditional solutions.
[0034] Fifth, the adaptive capability during the recycling phase is outstanding. Multiple GCL configurations achieve microsecond-level lossless switching through a double buffering mechanism, ensuring precise matching between communication resource allocation and the dynamic needs of the recycling phase. This avoids the resource waste problem caused by reserving all resources to adapt to the most demanding operating conditions in traditional solutions. Attached Figure Description
[0035] Figure 1 The diagram shown is a schematic representation of a motion cooperative control system based on machine learning and TSN deterministic communication in one embodiment of this application.
[0036] Figure 2 The diagram shown illustrates the specific implementation process of the upper-level recycling stage optimization of the joint optimization solver in one embodiment of this application.
[0037] Figure 3 The diagram shows a specific implementation process of the lower-level frame-level real-time scheduling of the joint optimization solver in one embodiment of this application.
[0038] Figure 4 The diagram shown is a flowchart illustrating a motion cooperative control method based on machine learning and TSN deterministic communication in one embodiment of this application.
[0039] Figure 5 The diagram shown is a structural schematic of an electronic terminal in one embodiment of this application. Detailed Implementation
[0040] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0041] In the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" do not necessarily imply that they are different.
[0042] It should be noted that, in the embodiments of this application, the words "exemplary" or "for example" indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0043] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0044] Before providing a further detailed description of the present invention, the nouns and terms used in the embodiments of the present invention are explained, and the nouns and terms used in the embodiments of the present invention are subject to the following interpretations:
[0045] <1> TSN (Time-Sensitive Networking) is a communication standard evolved from Ethernet that provides low-latency, low-jitter deterministic data transmission capabilities for industrial and control applications.
[0046] <2> GCL (Gate Control List): It is the core scheduling mechanism of the TSN network, used to define the transmission timing of each data frame and the rules for opening and closing ports to achieve orderly traffic transmission.
[0047] <3> MPC (Model Predictive Control): This algorithm relies on the system model to predict future states and solves for the optimal control quantity in a rolling manner. It is suitable for dynamic system control with time-series constraints.
[0048] <4> TT (Time-Triggered) Slot Ratio: Time-triggered, this indicator refers to the proportion of total communication slots occupied by time-triggered services, and is a core configuration parameter for TSN network resource allocation.
[0049] <5> AVB (Audio Video Bridging) Window Count: Audio and video bridging technology, the corresponding window count represents the number of dedicated transmission windows allocated in the network for audio and video data streams, used to ensure the quality of streaming media transmission.
[0050] <6> Gaussian Process (GP): A nonparametric probabilistic model often used as a surrogate model for function fitting, uncertainty assessment, and engineering optimization.
[0051] <7> Improved Expectation (EI) Acquisition Function: A commonly used acquisition function in Bayesian optimization, it comprehensively considers the needs of exploration and utilization, guides the algorithm to select the next sample point to be evaluated, and improves the optimization efficiency.
[0052] <8> BE (Best-Effort) window: This window is a transmission interval for non-real-time services in the TSN network, and no mandatory guarantee of latency or priority is made for such data.
[0053] <9> CBF (Control Barrier Function) safety filter: This filter constructs system safety constraints based on the control barrier function. It can correct control commands in real time, ensuring that the equipment's operating state is always limited to a safe range.
[0054] <10> NN-MPC (Neural Network Model Predictive Control) controller: This controller combines neural networks with model predictive control. It leverages neural networks to compensate for system model errors, improving the dynamic control accuracy and adaptability under complex operating conditions.
[0055] <11> Safety-oriented Reinforcement Learning (RL) Policy Networks: This network is the core carrier of reinforcement learning. It introduces safety constraints while learning the optimal control policy, balancing control performance and system operational safety.
[0056] <12> SQP (Sequential Quadratic Programming) algorithm: It is a mainstream algorithm for solving constrained nonlinear optimization problems and is widely used in scenarios involving control commands and optimal parameter solving.
[0057] To address the aforementioned technical problems, this invention provides a motion cooperative control system, method, terminal, and medium based on machine learning and TSN deterministic communication. The aim is to unify existing technical problems into a single communication-constrained real-time optimal control problem by constructing a TSN-AI joint optimization framework. This framework uses bounded end-to-end time delay as a coupling constraint and control performance, resource efficiency, and safety margin as joint optimization objectives. Through a two-level decomposition and online cooperative iterative solution strategy, it achieves joint optimality of communication and control with acceptable computational overhead.
[0058] It should be noted that the technical solution of this application can be applied to various motion devices with recycling requirements. Whether it's new energy vehicles and rail transit equipment in the transportation sector, lifting machinery and mobile robots in industrial settings, or even recyclable drones and manned return capsules in the aviation sector, all can be equipped with the deterministic communication, intelligent control, multi-actuator collaboration, and safety protection system of this solution. Based on core designs such as layered perception, adaptive mode switching, and dynamic allocation of control quantities, this solution can adapt to the motion characteristics and recycling conditions of different devices, completing status monitoring, trajectory control, fault redundancy, and mechanism linkage, effectively improving the stability, reliability, and control precision of the recycling process for various motion devices.
[0059] To facilitate understanding of the embodiments of this application, firstly, in conjunction with Figure 1 Detailed explanation. Figure 1 This illustration shows a schematic diagram of a motion cooperative control system based on machine learning and TSN deterministic communication in an embodiment of the present invention. The overall architecture of the motion cooperative control system in this embodiment includes the following four core layers: from top to bottom, they are the task planning layer, the joint optimization layer, the communication control execution layer, and the physical perception layer.
[0060] The task planning layer includes a task planner, which generates a full recycling task path based on the planning basis information and outputs prior information of the recycling stage to the joint optimization layer at a preset period, so that the parameter space of the joint optimization layer can be reduced in dimensionality according to the stage characteristics.
[0061] Preferably, the mission planner generates a full recovery mission path based on the target location, wind field prediction, and remaining fuel information. Generating the full recovery mission path includes calculating the reference trajectory of the motion device and defining the switching time nodes between each recovery sub-stage. The target location is a pre-set coordinate point where the motion device will finally dock, constraining the terminal position of the recovery mission and serving as the endpoint reference for the entire trajectory. Wind field prediction provides wind speed and direction forecasts at different altitudes and times across the entire mission area, used to offset the disturbances of atmospheric airflow on the motion device's attitude and trajectory. Remaining fuel information provides real-time statistics on the available propellant onboard the motion device, constraining the maximum power margin and range limit for the motion device's maneuverability.
[0062] Meanwhile, the task planner outputs the current recycling sub-stage identifier and task objective according to a preset cycle (e.g., a fixed refresh rate of 1Hz). This output is simultaneously transmitted to the joint optimization layer as prior information for the recycling stage. With the help of this prior information, the TSN-AI joint optimization engine in the joint optimization layer can compress the parameter range of the optimization solution based on the unique working characteristics of different recycling stages, achieving parameter space dimensionality reduction and significantly reducing the computational load of subsequent joint solutions.
[0063] The joint optimization layer includes a TSN-AI joint optimization engine, which comprises a network situation awareness module and a joint optimization solver. The network situation awareness module continuously collects actual end-to-end latency, network load status, and control performance indicators to construct a real-time network situation characterization. The joint optimization solver receives the network situation characterization and prior information from the recovery phase, and uses a two-layer decomposition strategy to solve for the optimal GCL configuration template and AI control parameters that achieve the joint objectives of control performance, resource efficiency, and security margin.
[0064] It should be noted that the actual end-to-end latency refers to the total time taken from the control terminal issuing control command data packets, through the TSN communication network to the actuator of the motion device, and then from the actuator to the control terminal sending back status messages. It is a core parameter for evaluating communication real-time performance. The network load status describes the current data transmission congestion level of the TSN communication link and switch ports, including indicators such as link bandwidth utilization, length of the waiting-to-forward buffer queue, and the proportion of data traffic for periodic and bursty services, which directly determine the probability of message queuing and congestion. The control performance indicators characterize the closed-loop control effect of the motion device. Available indicators include, but are not limited to, trajectory tracking error, attitude angle deviation, control overshoot, adjustment convergence time, and steady-state residual, reflecting the negative impact of communication latency fluctuations on motion control. After collecting these three types of heterogeneous data—actual end-to-end latency, network load status, and control performance indicators—they undergo preprocessing such as anomaly removal and dimension normalization. Then, a feature fusion algorithm is used to integrate the three isolated parameters into a network situational representation that changes in real time with the operating conditions.
[0065] The joint optimization solver employs a two-layer decomposition strategy, with the upper layer being a stage-level optimization for recycling and the lower layer being a frame-level real-time scheduling.
[0066] The upper-level recycling stage optimization process includes: running a Bayesian optimization algorithm at a preset cycle, modeling the joint objective function using a Gaussian process as a surrogate model, and searching for the optimal GCL configuration template and optimal AI controller parameters by improving the expected acquisition function. The specific implementation process of the upper-level recycling stage optimization is as follows: Figure 2 As shown, it includes the following steps:
[0067] Step S21: Perform dimensionality reduction of the high-dimensional parameter space based on prior information from the recycling phase.
[0068] The system receives prior information from the task planner for the recycling phase and, combined with GCL and MPC parameterization, compresses the high-dimensional communication and control parameters. Compression methods include, but are not limited to: replacing complete GCL entries with a few parameters such as TT slot percentage and AVB window number; replacing the massive MPC weight matrix with logarithmic values of the diagonal weight matrix and the prediction time-domain steps; and fixing the feasible range of some parameters through phase priors. Ultimately, the parameter space is compressed from hundreds of dimensions to 10-20 dimensions, reducing the computational burden for subsequent optimization.
[0069] It should be noted that GCL (Gate Control List) refers to the gated control list, and MPC (Model Predictive Control) refers to model predictive control. GCL parameterization is a simplified modeling method for Time-Sensitive Networks (TSNs). Instead of listing all gated control scheduling items one by one, it extracts a few feature variables that can constrain the operation characteristics of the entire GCL, using a small number of macroscopic parameters to represent the arrangement rules of the entire GCL, thus achieving dimensionality reduction of massive GCL items. MPC parameterization, on the other hand, condenses and abstracts the high-dimensional weight matrix inside the MPC controller, discarding all independent matrix elements in the weight matrix and using a few sets of key representative parameters to equivalently describe the controller tuning logic. GCL parameterization and MPC parameterization together achieve the convergence of communication parameters and control parameters from ultra-high dimension to low dimension.
[0070] TT (Time-Triggered) slot allocation refers to the percentage of time slot duration allocated to time-triggered services within a basic TSN communication cycle, used to constrain the bandwidth quota of periodic control messages. AVB (Audio Video Bridging) refers to audio and video bridging. The number of AVB windows represents the number of independent windows reserved in the scheduling cycle for transmitting audio and video, and non-hard real-time services. These two parameters are the core characteristic parameters used to summarize the overall gate control layout during GCL parameterization. A complete GCL entry is the original gate control list entry. GCL is essentially a gate opening and closing sequence table executed by switch ports in time sequence. One entry corresponds to the gate control action of one time slot, including the start and end times of the time slot, and detailed configurations such as opening or closing various queues such as TT / AVB / BE (Best-Effort). It should be understood that before parameterization, different flows correspond to dozens to hundreds of independent entries, and each configuration is considered an independent optimization variable. Therefore, the original GCL optimization dimension can reach hundreds of dimensions, which is the main reason for the huge parameter space.
[0071] The MPC weight matrix is a high-dimensional coefficient matrix that accompanies the cost function in the Model Predictive Control (MPC) algorithm. Each element in the matrix corresponds to a penalty coefficient for different state variables and control variables. The matrix size increases synchronously with the number of predicted time-domain variables and controlled states. Full matrix optimization will generate dozens or even hundreds of variables to be optimized. The prediction time-domain steps are the number of sampling steps that the MPC controller rolls forward to predict the future system state, determining the length of the controller's look-ahead calculation. The logarithmic value of the diagonal weight matrix is obtained by simplifying the original dense weight matrix into a diagonal matrix and then taking the logarithm of the diagonal coefficients. From the perspective of substitution logic, in aircraft recovery control scenarios, diagonal weights are mostly sufficient to meet control requirements. Off-diagonal elements are all set to zero and no longer participate in optimization. Only diagonal parameters are retained. Combined with logarithmic scaling to compress the parameter value range, and superimposed with the global configuration parameter of prediction time-domain steps, a few variables can achieve the equivalent control effect of the original entire high-dimensional weight matrix. This achieves the simplification and replacement of MPC parameters from high-dimensional to low-dimensional.
[0072] Step S22: Construct an agent model and optimization objective based on the real-time network situational representation.
[0073] The network situation is characterized in real time by obtaining network situation awareness module, with control performance, resource efficiency and security margin as joint optimization objectives; and Gaussian process is used as surrogate model to fit the mapping relationship between parameters and optimization objectives, thereby replacing the complex real solution process and greatly improving the optimization speed.
[0074] It should be understood that real-time network situational awareness includes three types of measured operating condition data: actual end-to-end latency, network load status, and control performance indicators. This data serves as the data source for quantifying the merits of the three joint optimization objectives and is also the input benchmark for Gaussian Process (GP) modeling. In this embodiment, the joint optimization objectives are: control performance refers to the aircraft's trajectory tracking accuracy, attitude control error, and command response lag level; resource efficiency represents the utilization rate of TSN network bandwidth and time slot resources; and safety margin refers to communication latency redundancy and power reserve margin, used to reserve fault tolerance space for sudden wind disturbances and instantaneous network congestion. It should be noted that there is a trade-off relationship among the three indicators. For example, improving resource utilization by compressing time slots can easily lead to excessive latency, deteriorated control, and reduced safety margin. Therefore, joint optimization seeks a balanced solution among the three to obtain the GCL and controller parameters that take into account flight reliability, resource conservation, and operational safety.
[0075] Gaussian processes, used as surrogate models, employ a GP mathematical model trained on samples to replace real physical simulations and airborne closed-loop experiments in calculating the optimization objective value. It should be understood that without a surrogate model, each adjustment of a set of GCL and MPC parameters requires sending data to the TSN deterministic communication subsystem and the AI recovery control subsystem for actual operational testing, collecting a round of network and control data before the objective function can be calculated. This is time-consuming in practice and cannot meet the upper-level optimization cycle of 100ms to 1s. By training the Gaussian process using existing network conditions and parameter samples, any subsequent set of candidate optimization parameters can be directly and quickly predicted using the GP model, thereby significantly reducing optimization time and improving optimization efficiency.
[0076] Step S23: Run Bayesian optimization to search for optimal parameters and output the optimal GCL configuration template and optimal AI control parameters.
[0077] Using the improved expectation (EI) acquisition function as the search criterion, the Bayesian optimization algorithm iteratively searches for the optimal parameter combination in the dimensionality-reduced parameter space to select the optimal parameter combination that enables the joint optimization objective. Finally, the optimized parameters are restored to a directly usable configuration form, and the finalized GCL configuration template and AI controller parameters are output.
[0078] It should be noted that Expected Improvement (EI) is the core sampling discriminant function in Bayesian optimization. Based on the mean and variance information output by the Gaussian process surrogate model, it simultaneously measures two search directions: first, the improvement in the target that candidate parameters can bring compared to the currently known optimal solution; and second, the model uncertainty corresponding to the parameter position. This function calculates the EI value for all candidate parameters in the parameter space. The higher the EI value, the more likely that the set of parameters can both optimize the overall objective and have high exploration value. The optimization algorithm prioritizes the parameters corresponding to the maximum EI value as the sampling points for the next round, serving as a quantitative benchmark for balancing local optimization and global exploration.
[0079] The reduced parameter space is a finite range of 10-20 dimensions for parameter values after GCL and MPC parameterization and stage prior constraints. Bayesian optimization iterative search is a closed-loop process of sampling, evaluating, and updating the surrogate model within this finite space. In each round, a set of candidate parameters is selected using the EI function, and the target value is predicted using a Gaussian model. If necessary, the actual joint target value is obtained from system measurements. Then, the Gaussian process surrogate model is refreshed with new parameter-target samples. The search range is narrowed repeatedly in multiple rounds to gradually approach the parameter point with the best overall performance. The optimized TT time slot ratio, AVB window number, and other feature parameters are obtained. The GCL gating entries with complete time-series arrangement are generated in reverse, forming a GCL configuration template with a fixed topology. For the MPC side, the complete diagonal MPC weight matrix is reconstructed from the optimized diagonal weight logarithmic value and the predicted time-domain steps, and all controller tuning parameters are completed. After the entire parameter conversion is completed, a standardized parameter set that can be directly distributed to the lower-level frame scheduling and airborne control system is generated.
[0080] The lower-level frame-level real-time scheduling process includes: performing frame-level deterministic scheduling at a preset frequency (e.g., 1kHz) based on the optimal GCL configuration template provided by the upper-level recycling stage optimization, to dynamically adjust the guard band size and micro-time slot boundaries without changing the GCL configuration template topology. It is worth noting that in actual system operation, network traffic, packet length, and device transmission latency will experience small dynamic fluctuations. Even with an optimal GCL configuration template, static time slots and intervals cannot adapt to instantaneous operating conditions. Therefore, it is necessary to fine-tune the guard band size and micro-time slot boundaries in real time, while maintaining the upper-level globally optimized scheduling logic and network architecture, and offsetting motion disturbances through frame-level dynamic correction.
[0081] The specific implementation process of lower-level frame-level real-time scheduling is as follows: Figure 3 As shown, it includes the following steps:
[0082] Step S31: Calculate the theoretical frame transmission time for each time-triggered time slot, obtain the measured jitter of the current time slot, and dynamically set the guard band width to the larger of the minimum guard band and twice the measured jitter.
[0083] Specifically, the TT time slots are traversed one by one. First, the theoretical transmission time of a single frame is calculated based on the message bytes and link rate. Then, the jitter value of the time slot transmission is captured from the actual communication measurement on site. The two values are compared: the fixed minimum guard band and twice the measured jitter. The larger one is selected as the new guard band width for this time slot, and the current time slot boundary position is initially modified.
[0084] Step S32: If setting the guard band width causes the start time of the next time slot to be delayed, then borrow time from the BE window; if the BE window has insufficient margin, then compress the elastic boundary of the AVB window. Specifically, updating the guard band width may cause the time slot occupancy time to exceed the limit and delay the start time of the next time slot. In this case, the surplus time slot duration of the BE (Best-Effort) window is used first to fill the gap; if the BE window has insufficient margin, then the available range of the AVB (AudioVideo Bridging) elastic window is compressed, and the timing offset is offset by the redundant duration of AVB. The overall topology of the GCL issued by the upper layer is not changed throughout the process, and only the boundaries of various windows and micro-time slots are fine-tuned.
[0085] Step S33: After the time slot boundary adjustment is completed, verify whether the new GCL still meets the latency upper bound constraint for all time-triggered flows. If the verification fails, restore the original configuration issued by the upper layer for stage-level optimization and send a re-optimization signal to the upper layer. Specifically, after all time slot boundaries are adjusted, verify whether the actual latency of all TT service flows is lower than the preset latency upper limit. If all links meet the constraints, the time slot configuration after this fine-tuning officially takes effect and is put into scheduling operation; if any time-triggered flow exceeds the latency index, immediately abandon this adjustment, restore the original GCL baseline configuration issued by the upper layer, and send a re-optimization request to the upper layer, which will iteratively generate a new GCL template.
[0086] The communication control execution layer includes the TSN deterministic communication subsystem and the AI recycling control subsystem. There is bidirectional data flow between the TSN-AI joint optimization engine and the downstream TSN deterministic communication subsystem and AI recycling control subsystem. That is, the TSN-AI joint optimization engine not only outputs the optimal configuration parameters to the downstream, but also receives performance feedback from the downstream to update the optimization model.
[0087] The TSN deterministic communication subsystem receives the optimal GCL configuration template and completes the data transmission of the entire system based on clock synchronization, lossless switching of multi-condition gating lists, and dual-link hot backup transmission, thus ensuring that the data transmission timing of the entire system is deterministic, the link is reliable, and the switching is uninterrupted.
[0088] Specifically, the TSN deterministic communication subsystem includes a gated list manager, a redundant dual-network switching matrix module, and a time synchronization module. The gated list manager and the time synchronization module are connected to the redundant dual-network switching matrix module. The gated list manager receives the optimal GCL configuration template from the TSN-AI joint optimization engine, manages multiple GCL templates corresponding to different recycling sub-stages, and performs microsecond-level lossless switching through a double-buffering mechanism. The redundant dual-network switching matrix module uses the IEEE 802.1CB frame duplication and elimination mechanism for dual-network hot backup to ensure that any single-network failure does not affect communication. The time synchronization module achieves sub-microsecond-level time synchronization of all moving nodes in the system based on the IEEE 802.1AS generalized precision time protocol (gPTP), and further improves synchronization accuracy by integrating GPS time sources and the short-term stability characteristics of inertial navigation.
[0089] It should be noted that the TSN deterministic communication subsystem is built on IEEE 802.1 (Institute of Electrical and Electronics Engineers). IEEE 802.1 is a suite of basic protocols for local area network bridging and time-sensitive networking developed by IEEE. The scheduling, redundancy, clock synchronization, and traffic control rules of the entire TSN deterministic communication system are all implemented according to this protocol standard.
[0090] The Gate Control List (GCL) manager is responsible for GCL configuration in the TSN subsystem. It continuously receives the optimal GCL configuration templates for each recovery stage from the TSN-AI joint optimization engine and stores multiple independent GCL schemes according to different recovery conditions of the motion device. The dual-buffering mechanism divides the configuration area into a running buffer and a standby buffer. The currently active GCL works in the running buffer, and the optimized configuration for the new stage is pre-written into the standby buffer. The buffer pointer is switched instantaneously at the stage switch, without interrupting message forwarding. This achieves microsecond-level uninterrupted and lossless configuration updates, ensuring that communication scheduling does not experience instantaneous disruptions when switching motion tasks.
[0091] The time synchronization module is the network-wide clock reference unit, which calibrates the local clocks of all boards, switches, controllers and other devices in the entire system to the same reference, achieving sub-microsecond high-precision clock synchronization.
[0092] The redundant dual-network switching matrix module first obtains the effective GCL time slot scheduling instructions from the gated list manager and the network-wide unified reference clock based on IEEE 802.1AS-gPTP (generalized Precision Time Protocol) from the time synchronization module. Based on the reference clock, it completes the time slot synchronization of all ports of the two physical switching networks A and B. Then, according to the service time slot opening and closing rules defined by GCL, it executes the IEEE 802.1CB (Frame Replication and Elimination for Reliability) dual-network redundancy mechanism at the same timing nodes. After the packet enters the forwarding port of the switching matrix, within the legal transmission time slot specified by GCL, the switching matrix generates two completely identical redundant packets for the same service frame according to the 802.1CB rules, and schedules them to the A network link and the B network link for synchronous forwarding. The receiving-side switching node, also under the constraints of unified clock and time slot, identifies the two packets from the same source carrying redundancy identifiers, retains the first arriving valid packet, and discards the delayed duplicate copies. Under normal circumstances, the dual-network parallel hot transmission achieves hot backup. When any one of the networks experiences a link break or a switch hardware failure that causes message interruption, the redundant messages from the other intact network can still reach the receiving end normally, and the receiving end can still receive valid data. This achieves redundant protection that does not interrupt communication in the event of a single network failure. The unified clock and GCL scheduling ensure that the transmission time slots of the two messages are strictly aligned, avoiding the failure of the 802.1CB deduplication logic caused by the timing disorder of the two network messages.
[0093] The AI-based recovery control subsystem generates control commands based on transmitted data and optimized parameters. It autonomously switches control modes according to operating conditions and performs safety corrections on the control commands, achieving intelligent closed-loop control and fault redundancy protection during the motion device recovery process. Specifically, the NN-MPC controller connects to the redundant dual-network switching matrix module in the TSN deterministic communication subsystem, and to the joint optimization solver in the TSN-AI joint optimization engine. It is also connected to the switching decision module and the CBF safety filter. The switching decision module is also connected to the safety RL policy network, which in turn is connected to the CBF safety filter.
[0094] Specifically, the AI recycling control subsystem includes an NN-MPC controller, a CBF safety filter, a safety RL policy network, and a switching decision module.
[0095] The NN-MPC controller is the main controller. It takes motion state estimation data as input, learns the residuals of the dynamic model through a neural network, and solves for the optimal control quantity in the rolling time domain (e.g., the NN-MPC controller performs rolling optimization every 2ms to 20ms). It should be understood that the NN-MPC controller loads the optimal AI control parameters as internal configuration to complete model and cost function tuning, and solves for the optimal control quantity in real time based on predetermined computational logic, improving dynamic control performance and adaptability to operating conditions. In NN-MPC, NN stands for Neural Network and MPC stands for Model Predictive Control. Motion state estimation data is collected and generated by inertial devices and various sensor units in the sensing and actuation layers of the motion device. The raw sensor messages are reliably transmitted via a redundant dual-network switching matrix module based on the TSN deterministic communication link, following GCL time slot scheduling and the 802.1CB dual-network redundancy mechanism, and finally delivered to the NN-MPC controller. It should be understood that the redundant dual-network switching matrix module relies on the synchronization clock module and the gating list to ensure that status messages are delivered on time with a predetermined delay, avoiding data out-of-order and delay jitter from interfering with the effectiveness of status estimation. On the other hand, it uses dual-network hot backup to avoid the interruption of status data transmission caused by single-link failure, ensuring the provision of continuous and complete motion status estimation data.
[0096] Specifically, the dynamic model of the NN-MPC controller adopts a residual learning structure, consisting of an analytical dynamic model and a neural network residual correction term. For example, the neural network residual here uses a three-layer fully connected structure with 128 and 64 neurons in the hidden layers, and employs INT8 quantization to meet the requirements of embedded real-time inference. This is for illustrative purposes only and is not a limitation.
[0097] Let the state vector of the motion device be... The control input vector is Then the state prediction formula is:
[0098] ;Formula (1)
[0099] in, For analytical dynamic model; This is a residual correction term for the neural network. These are the parameters of the neural network. The residual learning structure ensures that even if the neural network output is abnormal, the analytical dynamic model can still independently provide basic state prediction and control guarantees.
[0100] Within each control cycle, MPC solves the following finite-time optimal control problem:
[0101] ;Formula (2)
[0102] in, The variable to be optimized is represented by u, which is the control input vector, and the subscript is... This represents all control sequences to be solved for the next N consecutive steps starting from the current time t. `min` indicates searching for an optimal set of control sequences within the entire prediction time domain to minimize the total cost function. `N` is the number of prediction time domain steps, with a value between 10 and 30 that adaptively varies with the recovery phase. The reference trajectory is defined by: Q, the State Weight Matrix (positive definite diagonal matrix), used to quantify the penalty weights of different state components; a larger Q indicates a stricter constraint that the corresponding state cannot deviate from the reference value. R, the Input Weight Matrix (positive definite diagonal matrix), is used to constrain the magnitude of control variables and suppress frequent large movements of the actuator; a larger R results in greater fuel efficiency and limits drastic changes in driving force. P, the Terminal Weight Matrix (positive definite matrix), is used to constrain the terminal state at the last step t+N in the prediction time domain, forcing the predicted endpoint state to be close to the reference landing point.
[0103] It should be noted that the numerical solution for the optimal control quantity of this MPC uses the SQP (Sequential Quadratic Programming) algorithm. Conventional SQP iterations require numerical differentiation of the dynamic model of the motion device to obtain the Jacobian matrix. This application, however, uses an offline-trained neural network to directly output the Jacobian matrix to complete the system linearization, eliminating complex difference operations and accelerating the iteration convergence speed. At the same time, the system has a hard constraint that the time limit for a single optimization solution is 3 milliseconds to match the cycle time of the airborne control hardware. If a single SQP iteration exceeds the 3-millisecond time limit, to avoid control command lag and communication scheduling loss of synchronization, the controller directly uses the effective control quantity already solved in the previous control cycle as the current output, ensuring continuous and uninterrupted closed-loop control.
[0104] The CBF (Control Barrier Function) safety filter operates independently of the NN-MPC controller. It minimizes the modification of the original MPC control quantity without violating safety constraints, providing formal safety assurance. The CBF is a safety protection unit connected in series at the output of the NN-MPC controller, serving as a fallback correction for control commands. It should be understood that while the NN-MPC relies on SQP to solve for the baseline control quantity, the original control command may be affected by model errors and sudden airflow disturbances, potentially causing the motion device's flight state to exceed the physical safety boundaries of altitude, speed, and attitude. Therefore, the CBF safety filter uses quadratic programming to make small adjustments to the control quantity output by the MPC. While preserving the MPC's optimization intent to the greatest extent possible, it forcibly corrects the control command, ensuring that the motion state of the motion device always falls within the predefined safety set.
[0105] Specifically, the CBF security filter defines a security set. The safe set S consists of all functions that enable the barrier function. The established 12-dimensional motion state constitutes the structure. It includes three components: vertical velocity constraint, attitude angle constraint, and minimum height constraint. This represents a 12-dimensional real space, corresponding to a 12-dimensional full-state vector of the motion device (e.g., position, velocity, attitude angle, angular velocity, etc., totaling 12 state variables). The safety filtering process is achieved by solving the following quadratic programming problem:
[0106] ;Formula (3)
[0107] in, This represents the safety control quantity that is finally issued to the executing agency after CBF safety correction, which is the optimal solution to the quadratic programming optimization problem. Indicates in variable Within the range of values, find the control variable that minimizes the objective function; It is the squared objective term of the L2 norm. The original optimal control quantity is obtained by the NN-MPC controller through the SQP algorithm. The optimization objective is to make the corrected control quantity... The control parameters should be as close as possible to the original MPC control parameters, i.e., the extent of modification to the MPC control commands should be minimized. Constraints It is the core security inequality of CBF. It is the first derivative of the barrier function with respect to time. It is an extended class of K-functions, a nonlinear function whose values are strictly monotonically increasing and whose function value is zero at the origin. The entire process is a constrained quadratic programming problem, with a maximum solution time of 0.2 milliseconds per iteration. After the MPC output, a rapid safety correction is performed, and the final usable safety control command is output.
[0108] The safety RL (Reinforcement Learning) policy network acts as a backup controller for abnormal operating conditions, automatically taking over when the NN-MPC fails due to model mismatch or sensor degradation. Specifically, the safety RL policy network is the emergency backup control unit of the entire system, deployed in parallel with the main controller NN-MPC (Neural Network Model Predictive Control), but normally does not participate in the main control output, only continuously receiving motion state estimation data for online monitoring. When a model mismatch occurs, where the dynamic model of the moving device deviates significantly from the actual motion characteristics, or when sensor degradation failures occur due to decreased accuracy or abnormal data from the sensors on the moving device, the NN-MPC will be unable to solve for reliable control commands due to distorted input data and model prediction failure, triggering fault judgment logic. At this time, the safety RL policy network immediately and automatically switches to the working state, taking full control authority. Based on the fault-tolerant control strategy pre-trained through reinforcement learning, it outputs safe control quantities to ensure that the moving device does not deviate from the safe zone due to the failure of the main controller, thereby achieving system redundancy and fault tolerance under abnormal operating conditions.
[0109] The switching decision module switches between three control modes—NN-MPC dominant mode, parallel fusion mode, and safe RL independent mode—based on the neural network residual prediction error. Specifically, the switching strategy includes: if the duration of the neural network residual prediction error exceeding a preset threshold does not exceed a preset alarm period, the NN-MPC dominant mode is maintained; if the duration of the neural network residual prediction error exceeding the preset threshold continuously exceeds the preset alarm period, the parallel fusion mode is switched, and the safe RL independent mode is switched if the NN-MPC controller cannot solve the problem normally.
[0110] It should be noted that the NN-MPC-dominated mode refers to the NN-MPC controller acting as the main controller to output optimized control commands under normal, fault-free operating conditions. The CBF safety filter only performs fine-tuning of safety boundaries, and the safety RL policy network remains in the background without participating in the output. The parallel fusion mode refers to the NN-MPC controller and the safety RL policy network synchronously outputting control quantities and then weighting and fusing them. The safety RL independent mode means that the safety RL policy network takes full control.
[0111] Specifically, the formula for calculating the prediction error of the neural network residual model in each control cycle by the switching decision module is as follows:
[0112] ;Formula (4)
[0113] in, Indicates the actual measurement obtained by the sensor The real flight status of the motion device at all times; This represents a physical dynamics model based on the mechanism of motion devices, used to rely on Moment State With control quantity Perform basic state simulation; It is a neural network residual model used by the NN-MPC controller to compensate for mechanistic biases. This represents the weight parameters of the neural network; the result is obtained by taking the 2-norm after subtracting the three terms. , The larger the value, the more serious the deviation between the model and the actual motion characteristics, and the higher the risk of NN-MPC modeling failure.
[0114] In addition, set an adaptive threshold. ,in and The mean and standard deviation within the sliding window. These are sub-stage related correction terms. When... If the alarm cycle count continues to exceed the preset number, a switch will be triggered from NN-MPC dominant mode to non-integrated mode or security RL independent mode.
[0115] The multi-actuator collaborative subsystem includes a control distributor and multiple actuators. The control distributor is connected to the CBF safety filter in the AI recovery control subsystem and is electrically connected to each actuator. The control distributor of the multi-actuator collaborative subsystem receives the optimal control quantity output from the CBF safety filter and decomposes it into individual execution control commands through a weighted pseudo-inverse allocation method, which are then sent to the servo control loops of each actuator. It should be noted that the weighted pseudo-inverse allocation adds a diagonal weight matrix to the basic Moore-Penrose pseudo-inverse, decomposing the total virtual control command output from the upper layer into actual drive commands for each actuator such as the control surface and thruster, taking into account optimal allocation, energy consumption, and mechanism priority constraints.
[0116] The physical sensing layer includes a sensor group and an actuator group, which are connected to the redundant dual-network switching matrix module in the TSN deterministic communication subsystem.
[0117] The sensor group includes, but is not limited to, IMU (Inertial Measurement Unit), lidar, radar altimeter, GPS (Global Positioning System), engine condition sensor, landing leg position sensor, etc. Various devices collect raw sensing data of motion status in real time. This data is processed through a redundant dual-network switching matrix, scheduled and transmitted according to GCL gated time slots, and then summarized and filtered to form a motion state estimate, which serves as the input data source for the NN-MPC controller.
[0118] The actuator group includes, but is not limited to, engine thrust vectoring mechanisms, grid rudder servo mechanisms, and landing leg deployment and locking mechanisms. Specifically, the multi-actuator coordination subsystem sends execution control commands to the servo control loops of each actuator. Each servo control loop then drives its corresponding servo mechanism to perform the corresponding operation. During the execution process, the actuator group collects latency, control error, and safety margin through performance monitoring, and then feeds the performance back to the joint optimization solver of the TSN-AI joint optimization engine through the TSN deterministic communication subsystem, thereby continuously updating the optimization model.
[0119] In a specific embodiment: The present invention relies on a co-simulation platform to complete performance verification. This platform consists of a dynamics simulator, a TSN network simulator, and an AI controller simulator. Each subsystem achieves time synchronization and coupled computation with a step size of 0.1 milliseconds through the co-simulation engine. The dynamics simulator uses a six-degree-of-freedom model to build aerodynamic, dynamic, and wind disturbance models to reproduce the dynamic characteristics of the recovery process; the TSN network simulator is based on the OMNeT+INET framework to simulate network mechanisms such as gating list scheduling, frame preemption, frame duplication, and elimination; the AI controller simulator is based on the PyTorch framework to complete the computational inference of NN-MPC and safety RL policy networks.
[0120] This simulation included several typical scenarios, such as standard operation without interference, crosswind, gust wind shear, engine thrust decay, and network failure. Under standard operation, the expected circular error of the landing point was no greater than 3 meters, and the landing vertical velocity was no higher than 1.5 m / s. The end-to-end control delay of the final approach segment was capped at 2 milliseconds, and the standard deviation of the delay jitter was no more than 10 microseconds. When engine thrust decreased by 20%, the system could switch to a safe independent RL control mode within 5 milliseconds to ensure operational safety.
[0121] After quantization using the FPGA platform's INT8, the residual network inference time is no more than 0.5 milliseconds, and the overall NN-MPC solution time is no more than 3 milliseconds, which can adapt to the hard real-time requirement of a minimum control cycle of 2 milliseconds. If the solution times out, the safety control strategy of the previous cycle is used. The safety RL strategy network inference time is also no more than 0.5 milliseconds, and the GCL double-buffer switching time is no more than 1 microsecond. The above indicators are all design and simulation expectations. The final performance needs to be further verified by combining high-fidelity simulation and ground tests.
[0122] It should be understood that the specific processes by which each module performs the corresponding steps described above have been detailed in the above method embodiments, and will not be repeated here for the sake of brevity. It should also be understood that the module division in the embodiments of this application is illustrative and merely a logical functional division; other division methods may exist in actual implementation. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or have two or more modules integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0123] It is also worth emphasizing that the core difference of this invention lies in treating TSN communication resource configuration and AI control strategy as two coupled dimensions of a unified system, solving them jointly within the same optimization framework. Specifically, traditional solutions employ a hierarchical decoupling model, designing the communication network first and then the control strategy. Communication resource configuration is independent of control requirements, and the control strategy assumes fixed communication performance. This invention, however, establishes a joint communication-control optimization model, using the upper bound of end-to-end communication delay not exceeding the control cycle as a coupling constraint. Through a two-layer optimization architecture, it online searches for the optimal gating list configuration and controller parameter combination, enabling dynamic allocation of communication resources according to control requirements and adapting the control strategy to actual communication capabilities, achieving Pareto optimality for global performance.
[0124] Figure 4 This application illustrates a flowchart of a motion cooperative control method based on machine learning and TSN deterministic communication, according to an embodiment of the present application. The motion cooperative control method includes:
[0125] Step S41: Invoke the task planner to generate a full recycling task path based on the planning basis information, and output the prior information of the recycling stage to the joint optimization layer at a preset period.
[0126] Step S42: Mobilize the TSN-AI joint optimization engine to construct a real-time network situational representation, and based on the network situational representation and prior information from the recovery phase, set a joint objective and use a two-layer decomposition strategy to solve for the optimal GCL configuration template and the optimal AI control parameters.
[0127] Step S43: The TSN deterministic communication subsystem is activated to receive the optimal GCL configuration template, and the data transmission of the entire system is completed based on clock synchronization, lossless switching of multi-condition gating lists and dual-link hot backup transmission; and the AI recovery control subsystem is activated to generate the optimal control quantity based on the optimal AI control parameters and real-time motion state estimation; the multi-actuator coordination subsystem receives the optimal control quantity and generates the corresponding servo control command.
[0128] Step S44: Activate the sensor group to collect raw sensing data of motion state, and use the actuator group to perform corresponding operations.
[0129] It should be noted that the implementation process and principle of the motion cooperative control method based on machine learning and TSN deterministic communication provided in this application embodiment are similar to those of the motion cooperative control system based on machine learning and TSN deterministic communication described above, and will not be repeated here.
[0130] Figure 5 This is a schematic block diagram of the electronic terminal provided in the embodiments of this application. Figure 5 As shown, the electronic terminal 500 includes at least one processor 501, a memory 502, at least one network interface 503, and a user interface 505. The various components in the electronic terminal 500 are coupled together via a bus system 504. It is understood that the bus system 504 is used to implement communication between these components. In addition to a data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 5 The general will label all buses as bus systems.
[0131] The user interface 505 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.
[0132] It is understood that memory 502 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.
[0133] In this embodiment of the invention, the memory 502 is used to store various types of data to support the operation of the electronic terminal 500. Examples of this data include: any executable program for operation on the electronic terminal 500, such as the operating system 5021 and application programs 5022; the operating system 5021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 5022 may contain various applications, such as a media player, browser, etc., for implementing various application services. The methods provided in this embodiment of the invention can be included in the application program 5022.
[0134] The methods disclosed in the above embodiments of the present invention can be applied to processor 501, or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 501 or by instructions in the form of software. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 501 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 501 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.
[0135] In an exemplary embodiment, the electronic terminal 500 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the aforementioned method.
[0136] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to perform the method of any of the embodiments described above.
[0137] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to perform the method of any of the embodiments described above.
[0138] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0139] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0140] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0141] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0142] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0143] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0144] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0145] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0146] In summary, this application provides a motion cooperative control system, method, terminal, and medium based on machine learning and TSN deterministic communication. This application effectively overcomes the various shortcomings of the prior art and has high industrial application value.
[0147] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A motion cooperative control system based on machine learning and TSN deterministic communication, characterized in that, include: The task planning layer includes the task planner; The task planner generates a full recovery task path based on the planning basis information and outputs prior information of the recovery stage to the joint optimization layer at a preset period, including: the task planner generates a full recovery task path based on the target location point, wind field prediction and fuel remaining information. The process of generating the full recovery task path includes calculating the reference motion trajectory of the motion device and defining the switching time nodes between each recovery sub-stage; the prior information of the recovery stage includes the current recovery sub-stage identifier and the task objective. The joint optimization layer includes a TSN-AI joint optimization engine; the TSN-AI joint optimization engine is connected to the task planner and includes a network situation awareness module and a joint optimization solver; the network situation awareness module is used to construct a real-time network situation representation; the joint optimization solver receives the network situation representation and prior information of the recovery phase, sets a joint objective, and uses a two-layer decomposition strategy to solve for the optimal GCL configuration template and the optimal AI control parameters. The communication control execution layer includes a TSN deterministic communication subsystem, an AI recovery control subsystem, and a multi-actuator coordination subsystem. The TSN deterministic communication subsystem receives the optimal GCL configuration template and completes data transmission for the entire system based on clock synchronization, lossless switching of multi-condition gating lists, and dual-link hot backup transmission. The AI recovery control subsystem generates the optimal control quantity based on the optimal AI control parameters and real-time motion state estimation. The multi-actuator coordination subsystem receives the optimal control quantity and generates corresponding servo control commands. The physical sensing layer includes a sensor group and an actuator group; both the sensor group and the actuator group transmit data through the TSN deterministic subsystem.
2. The motion cooperative control system based on machine learning and TSN deterministic communication according to claim 1, characterized in that, The network situation awareness module continuously collects actual end-to-end latency, network load status, and control performance indicators to construct a real-time network situation representation.
3. The motion cooperative control system based on machine learning and TSN deterministic communication according to claim 1, characterized in that, The joint optimization solver employs a two-layer decomposition strategy, including: upper-layer stage-level optimization and lower-layer frame-level real-time scheduling, wherein: The upper-level recycling stage optimization process includes: running the Bayesian optimization algorithm at a preset cycle, modeling the joint objective function using a Gaussian process as a surrogate model, and searching for the optimal GCL configuration template and optimal AI controller parameters by improving the expected acquisition function; The lower-level frame-level real-time scheduling process includes: performing frame-level deterministic scheduling at a preset frequency based on the optimal GCL configuration template provided by the upper-level recycling stage optimization, so as to dynamically adjust the guard band size and micro-slot boundary without changing the GCL configuration template topology.
4. The motion cooperative control system based on machine learning and TSN deterministic communication according to claim 3, characterized in that, The process of running a Bayesian optimization algorithm at a preset period and modeling the joint objective function using a Gaussian process as a surrogate model, and searching for the optimal GCL configuration template and optimal AI controller parameters by improving the expected acquisition function, includes: Based on the prior information of the recycling phase, and combined with GCL parameterization and MPC parameterization, the high-dimensional communication and control parameters are compressed. Based on network situational characterization, with control performance, resource efficiency, and security margin as joint optimization objectives, and using a Gaussian process as a surrogate model, the mapping relationship between parameters and optimization objectives is fitted. Using the improved expected acquisition function as the search criterion, the GCL configuration template and AI controller parameters that make the joint optimization objective optimal are selected through iterative optimization in the dimensionality-reduced parameter space by Bayesian optimization algorithm.
5. The motion cooperative control system based on machine learning and TSN deterministic communication according to claim 1, characterized in that, The TSN deterministic communication subsystem includes a gated list manager, a redundant dual-network switching matrix module, and a time synchronization module; the gated list manager and the time synchronization module are respectively connected to the redundant dual-network switching matrix module, and the gated list manager receives the optimal GCL configuration template.
6. The motion cooperative control system based on machine learning and TSN deterministic communication according to claim 1, characterized in that, The AI recovery control subsystem includes an NN-MPC controller, a CBF safety filter, a safety RL policy network, and a switching decision module. The NN-MPC controller is connected to the CBF safety filter and the switching decision module, and is also connected to the joint optimization solver to obtain the optimal AI control parameters. The safety RL policy network is also connected to the CBF safety filter and the switching decision module. Wherein: The NN-MPC controller takes motion state estimation data as input, learns the dynamic model residuals through a neural network, and solves the optimal control quantity in the rolling time domain. The CBF safety filter is used to correct the optimal control quantity output by the NN-MPC controller; The secure RL policy network is used to automatically take over when the NN-MPC controller fails; The switching decision module switches between three control modes—NN-MPC dominant mode, parallel fusion mode, and safety RL independent mode—based on the neural network residual prediction error. In the parallel fusion mode, the NN-MPC controller and the safety RL policy network synchronously output control quantities and then weighted and fused them.
7. The motion cooperative control system based on machine learning and TSN deterministic communication according to claim 6, characterized in that, The switching strategy of the switching decision module includes: if the duration of the neural network residual prediction error exceeding the preset threshold does not exceed the preset alarm period, then the NN-MPC dominant mode is maintained; if the duration of the neural network residual prediction error exceeding the preset threshold continuously exceeds the preset alarm period, then the parallel fusion mode is switched, and the safe RL independent mode is switched when the NN-MPC controller cannot solve normally.
8. A motion cooperative control method based on machine learning and TSN deterministic communication, characterized in that, The motion cooperative control method is applied to the motion cooperative control system based on machine learning and TSN deterministic communication as described in claim 1; the motion cooperative control method includes: The task planner is mobilized to generate a full recycling task path based on the planning basis information, and the recycling phase prior information is output to the joint optimization layer at a preset period. The TSN-AI joint optimization engine is mobilized to construct a real-time network situational representation. Based on the network situational representation and prior information from the recovery phase, a joint objective is set and a two-layer decomposition strategy is adopted to solve for the optimal GCL configuration template and the optimal AI control parameters. The TSN deterministic communication subsystem is activated to receive the optimal GCL configuration template, and data transmission of the entire system is completed based on clock synchronization, lossless switching of multi-condition gating lists, and dual-link hot backup transmission; and the AI recovery control subsystem is activated to generate the optimal control quantity based on the optimal AI control parameters and real-time motion state estimation; the multi-actuator coordination subsystem receives the optimal control quantity and generates corresponding servo control commands. The sensor array is activated to collect raw sensor data on motion state, and the actuator array is used to perform corresponding operations.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the motion cooperative control method based on machine learning and TSN deterministic communication as described in claim 8.
10. An electronic terminal, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the motion cooperative control method based on machine learning and TSN deterministic communication as described in claim 8.
Citation Information
Patent Citations
Carrier rocket avionics system based on TSN time sensitive network and carrier rocket avionics data transmission method
CN122069238A
Power 5G low-delay jitter implementation method for distribution network stability protection
CN114827195A
Industrial production line multi-equipment dynamic collaborative scheduling method and system based on reinforcement learning
CN120630910A